gatk-3.8

Commit Graph

Author	SHA1	Message	Date
Eric Banks	7e0c326e1c	Merge pull request #706 from broadinstitute/vrr_reduce_hc_integration_test_time Reduce intervals of integration tests in HaplotypeCallerIntegrationTest ...	2014-08-15 17:37:57 -04:00
Valentin Ruano-Rubio	2f79042dee	Reduce intervals of integration tests in HaplotypeCallerIntegrationTest class Story: https://www.pivotaltracker.com/story/show/74858854 Changes: Intervals have been shrunk so that the test run in 15s or less.	2014-08-15 14:20:10 -04:00
Eric Banks	eb84091702	Update the --keepOriginalAC functionality in SelectVariants to work for sites that lose alleles in the selection.	2014-08-14 15:34:09 -04:00
Ryan Poplin	3a9a78c785	Removing an assumption that ADs were in the same order if the number of alleles matched. This happens for example when one sample is C->T and another sample is C->G.	2014-08-13 13:26:40 -04:00
Eric Banks	27193c5048	Merge pull request #700 from broadinstitute/eb_phase_HC_variants_PT74816060 Initial implementation of functionality to add physical phasing informat...	2014-08-13 12:30:32 -04:00
Eric Banks	4512940e87	Initial implementation of functionality to add physical phasing information to the output of the HaplotypeCaller. If any pair of variants occurs on all used haplotypes together, then we propagate that information into the gVCF. Can be enabled with the --tryPhysicalPhasing argument.	2014-08-13 12:25:31 -04:00
Valentin Ruano-Rubio	b39508cd15	ReadLikelihoods class introduction final changes before merging Stories: https://www.pivotaltracker.com/story/show/70222086 https://www.pivotaltracker.com/story/show/67961652 Changes: Done some changes that I missed in relation with making sure that all PairHMM implentations use the same interface; as a consequence we were running always the standard PairHMM. Fixed some additional bugs detected when running it on full wgs single sample and exom multi sample data set. Updated some integration test md5s.	2014-08-11 17:47:25 -04:00
Valentin Ruano-Rubio	9a9a68409e	ReadLikelihoods class introduction final changes before merging Stories: https://www.pivotaltracker.com/story/show/70222086 https://www.pivotaltracker.com/story/show/67961652 Changes: Done some changes that I missed in relation with making sure that all PairHMM implentations use the same interface; as a consequence we were running always the standard PairHMM. Fixed some additional bugs detected when running it on full wgs single sample and exom multi sample data set. Updated some integration test md5s. Fixing GraphBased bugs with new master code Fixed ReadLikelihoods.changeReads difficult to spot bug. Changed PairHMM interface to fix a bug Fixed missing changes for various PairHMM implementations to get them to use the new structure. Fixed various bugs only detectable when running with full sample(s). Believe to have fixed the lack of annotations in UG runs Fixed integrationt test MD5s Updating some md5s Fixed yet another md5 probably left out by mistake	2014-08-11 17:46:28 -04:00
Valentin Ruano-Rubio	0b472f6bff	Added new test to verify the functionality of ReadLikelihoods.java and its use in HC. Updated existing integration test md5s. Stories: https://www.pivotaltracker.com/story/show/70222086 https://www.pivotaltracker.com/story/show/67961652	2014-08-11 17:46:28 -04:00
Valentin Ruano-Rubio	2914ecb585	Change the Map-of-maps-of-maps for an array based implementation ReadLikelihoods to hold read likelihoods. The array structure should be faster to populate and query (no properly benchmarked) and reduce memory footprint considerably. Nevertheless removing PairHMM factor (using likelihoodEngine Random) it only achieves a speed up of 15% in some example WGS dataset i.e. there are other bigger bottle necks in the system. Bamboo tests also seem to run significantly faster with this change. Stories: https://www.pivotaltracker.com/story/show/70222086 https://www.pivotaltracker.com/story/show/67961652 Changes: - ReadLikelihoods added to substitute Map<String,PerSampleReadLikelihoods> - Operation that involve changes in full sets of ReadLikelihoods have been moved into that class. - Simplified a bit the code that handles the downsampling of reads based on contamination Caveats: - Still we keep Map<String,PerReadAlleleLikelihoodsMap> around to pass to annotators..., didn't feel like change the interface of so many public classes in this pull-request.	2014-08-11 17:46:28 -04:00
Ryan Poplin	c56e493f98	Merge pull request #622 from broadinstitute/ldg_SORanalysis Add StrandOddsRatio to default annotations produced by GenotypeGVCFs	2014-08-11 09:45:27 -04:00
Tim Fennell	5695f22da8	Changed the default GVCF Q Bands from 5,20,60 to be 1..60 by 1s, 60...90 by 10s and 99 in order to give finer resolution for homref PLs and ADs at lower confidences and somewhat higher resolution at higher confidences.	2014-08-08 14:31:35 -04:00
Laura Gauthier	35de598e4b	Modify StrandOddsRatio calculation to take on lower values in cases where reference +/- reads are skewed but alt reads are not. Add SOR to default annotations produced by GenotypeGVCFs. Add jitter to minimum SOR values	2014-08-07 12:09:19 -04:00
Laura Gauthier	f532f1f843	Fix nullPointerException	2014-08-07 10:13:17 -04:00
Laura Gauthier	74affcc077	Update inbreeding coefficient calculation to give a better estimate for multialleleic sites Add unit test for compound het and for multiallelic hets	2014-08-07 08:12:47 -04:00
Eric Banks	b9486f5b4d	Merge pull request #693 from broadinstitute/ldg_SORfromHC Allow SOR to be calculated from HC	2014-08-06 21:48:09 -04:00
Phillip Dexheimer	593663d9b6	Improved detection of missing argument values In particular, it was possible to specify arguments for Files or Compound types without values Added a special "none" value for annotations, since a bare "-A" is no longer allowed Delivers PT 71792842 and 59360374	2014-08-05 20:31:31 -04:00
Laura Gauthier	5533199402	Allow SOR to be calculated from HC Refactor StrandBiasTest classes	2014-08-01 20:47:58 -04:00
Ryan Poplin	63b3f7dfd3	Fixing typos in AnalyzeCovariates	2014-07-31 10:36:18 -04:00
Valentin Ruano-Rubio	750eb4b5a6	Add diploid only support message to HaplotypeCaller Story: https://www.pivotaltracker.com/story/show/73440292 Changes: - Just add the conditional in HaplotypeCaller#initialize Testing: - Nothing added, checked locally, trivial change that would eventually be removed anyway.	2014-07-29 17:05:36 -04:00
David Roazen	0798a4b768	Update pom versions to mark the start of GATK 3.3 development	2014-07-17 12:09:33 -04:00
David Roazen	323f22f852	Update pom versions for the 3.2 release	2014-07-17 12:06:22 -04:00
Eric Banks	98d88eb07e	Fixed IndexOutOfBounds error associated with tail merging. Don't expand out source nodes for tail merging, since that's a head merging action only. This shows up as a bug only because we now allow merging tails against non-reference paths.	2014-07-17 12:04:22 -04:00
Geraldine Van der Auwera	a6f632874b	Various documentation improvements - Edited intervals merging docs for correctness & clarity - Edited VQSR arg docs and made mode required (+added -mode SNP to VQSR tests) - Moved PaperGenotyper to Toy Walkers to declutter the actually useful docs - Moved GenotypeGVCFs to Variant Discovery category and clarified a few points - Clarified that the -resource argument depends on using the -V:tag format - Clarified how the pcr indel model works - Added caveat for -U ALLOW_N_CIGAR_READS - Added MathJax support for displaying equations in GATKDocs - Updated HC example commands and caveats	2014-07-14 12:03:03 -04:00
droazen	db53d096c9	Merge pull request #684 from broadinstitute/ks_add_cofoja_to_gatk_packages Added cofoja to the gatk packages for tests to pass.	2014-07-14 11:15:49 -04:00
Eric Banks	ecefcb383d	Disable the complex variant merging for now, as requested by ATGU	2014-07-11 17:27:40 -04:00
Khalid Shakir	c7e357eb59	Added cofoja to the gatk packages for tests to pass.	2014-07-11 23:19:42 +08:00
droazen	b8751ad598	Merge pull request #680 from broadinstitute/ldg_VQSRscript Update VQSR Rnd BQSR script generation code for compatibility with late...	2014-07-11 10:16:37 -04:00
Eric Banks	1d97b4a191	Improved tail merging: now tails can be merged to branches that are not entirely reference. This is useful for e.g. cases where there are SNPs on insertions. Before tails were forced to be merged (incorrectly) only to a reference node, but now they can be merged to any path in the graph from which they directly branch. Also, I've transferred over Ryan's code to refuse to process kmer sizes such that there are non-unique kmers in the reference sequence with them.	2014-07-10 08:57:01 -04:00
Ryan Poplin	5eee065133	Merge pull request #674 from broadinstitute/rp_improve_genotyping Improvements to genotyping accuracy.	2014-07-09 16:03:09 -04:00
Laura Gauthier	99026eb51b	Update VQSR Rnd BQSR script generation code for compatibility with latest ggplot version. Update queueJobReport.R and public/gsalib/src/R/R/gsa.variantqc.utils.R also	2014-07-09 15:36:58 -04:00
Ryan Poplin	74a7674d70	Improvements to genotyping accuracy. -- Global mismapping penalty was only applied to the reference haplotype. This led to problems with overlapping events, mostly STR haplotypes. Now the penalty is applied to every haplotype. -- We subset the reads down to only those which overlap the event (after assembly based realignment) for likelihood calculations.	2014-07-09 13:11:07 -04:00
David Roazen	719e685759	Remove junit imports in the test suite	2014-07-09 12:09:27 -04:00
Eric Banks	bad7865078	When converting a haplotype to a set of variants we now check for cases that are overly complex. In these cases, where the alignment contains multiple indels, we output a single complex variant instead of the multiple partial indels. We also re-enable dangling tail recovery by default.	2014-07-01 14:18:59 -04:00
Ryan Poplin	e14bff212d	SB tables should be created even if the ref or alt columns have no counts. This is so that FS/SOR will still be calculated when the variant is extremely high or low frequency. -- Removed long running HC integration test... sorry	2014-06-30 15:19:15 -04:00
Ryan Poplin	0127799cba	Reads are now realigned to the most likely haplotype before being used by the annotations. -- AD,DP will now correspond directly to the reads that were used to construct the PLs -- RankSumTests, etc. will use the bases from the realigned reads instead of the original alignments -- There is now no additional runtime cost to realign the reads when using bamout or GVCF mode -- bamout mode no longer sets the mapping quality to zero for uninformative reads, instead the read will not be given an HC tag	2014-06-30 10:35:50 -04:00
Phillip Dexheimer	06d619e9aa	Removed redundant SelectVariantsIntegrationTest, merged it's only test into protected version	2014-06-24 18:59:59 -04:00
Eric Banks	2df2a153e6	Merge pull request #658 from broadinstitute/ldg_PbyTwithPriors Updated CalculateGenotypePosteriors to compute genotype posteriors using...	2014-06-18 15:04:39 -04:00
Laura Gauthier	2356d5d63f	Updated CalculateGenotypePosteriors to compute genotype posteriors using likelihoods from all members of the trio. (Right now it only works if all members of the trio are called.) Takes posteriors as input, defaulting to PLs Added annotations for possible de novos for us in full genotype refinement pipeline Added family priors to CGP integration test. Changed CGP to use PP tag instead of GP tag because posteriors are Phred-scaled. Updated CGP integration test md5s to reflect change.	2014-06-18 11:17:15 -04:00
Phillip Dexheimer	2e78815055	Added missing arguments to GenotypeGVCFs - New arguments are nda, hets, indelHeterozygosity, stand_call_conf, stand_emit_conf, ploidy, and maxAltAlleles - Addresses PT 70110918 - To do this, moved those arguments out of the StandardCallerArgumentCollection into a new GenotypeCalculationArgumentCollection, which is now included as a member of SCAC	2014-06-16 08:10:54 -04:00
droazen	3079755b4c	Merge pull request #646 from broadinstitute/ks_disable_distribution_with_private Add maven -Pgsadev flag to build private jars only	2014-06-11 11:00:31 -04:00
Khalid Shakir	f082572593	If passed -Pgsadev, don't build the distribution package.	2014-06-10 23:33:33 -04:00
Valentin Ruano Rubio	db96891d4b	Merge pull request #638 from broadinstitute/vrr_createTempFile_testfix Changed File.createTempFile to BaseTest.createTempFile calls Test	2014-05-29 10:15:05 -04:00
Valentin Ruano-Rubio	07567fdae3	Removed debug code outputing files not removed after VM exists in ReadThreadingLikelihoodCalculationEngineUnitTest. Notice however that this should not be the cause of resent problems as the code was desactivated.	2014-05-28 19:03:25 -04:00
Valentin Ruano-Rubio	e0c221470c	Changed File.createTempFile to BaseTest.createTempFile	2014-05-28 18:59:48 -04:00
EvolvedMicrobe	ef7531d4a5	Merge pull request #640 from broadinstitute/IntegerSWImplementation Change SmithWaterman to use integers instead of doubles.	2014-05-28 15:10:05 -04:00
Nigel Delaney	cc45e62e8e	Change SmithWaterman to use integers instead of doubles.	2014-05-28 13:13:14 -04:00
droazen	ac52fa581a	Merge pull request #644 from broadinstitute/ks_queue_test_temp_fix Disabled ExampleUG Queue Tests, fixed internal extensions dependency.	2014-05-28 11:29:08 -04:00
Phillip Dexheimer	c15e6fcc0e	Refactored the static lookup arrays in MathUtils (log10Cache, log10FactorialCache, jacobianLogTable) -They are now only computed when necessary -Log10Cache is dynamically resizable, either by calling get() on an out-of-range value or by calling ensureCacheContains -Log10FactorialCache and JacobianLogTable are initialized to a fixed size on first access and are not resizable -Addresses PT 69124396	2014-05-27 22:27:57 -04:00
Eric Banks	b77589696e	Merge pull request #643 from broadinstitute/rp_remove_hwp Removing HWP from GenotypeSummaries because of integer overflow issues w...	2014-05-27 17:21:19 -04:00

1 2 3 4 5 ...

1146 Commits (7e0c326e1c02e34fda720e8c0d3f00256f6a2781)