gatk-3.8

Commit Graph

Author	SHA1	Message	Date
Valentin Ruano-Rubio	95b45443ae	Updated test according to changes in the AF calculator framework. Changes: ------- * Updated current unit and integration test to use the new API components. * Added unit tests for new classes AFPriorProvider and AFCalculatorProviders. * Added integration test for mixed ploidy GenotypeGVCFs and CombineGVCFs	2014-09-12 14:59:47 -04:00
Valentin Ruano-Rubio	3cdeab6e9e	GenotypingEngines and walkers now use AFCalc(ulator) providers rathern than instanciate their own (fixed) calculators directly. Changes: ------- * GenotypingEngine uses now a AFCalc provider instead of its own thread-local with one-time initialized and fixed AF calculator. * All walkers that use a GenotypingEngine now are passing the appropiate AF calculator provider. For now most just use a fix calculator (FixedAFCalculatorProvider) except GenotypeGVCFs as this one now can cope with mixture of ploidies failing-over to a general-ploidy calculator when the preferred implementation is not capable to handle a site's analysis.	2014-09-12 14:25:09 -04:00
Valentin Ruano-Rubio	935bd1394b	AFCalculatorProvider components to allow for dynamic instantiation of different AFCalc(ulators) to cope with dynamic ploidy and max-alt-allele counts (the latter not used for now).	2014-09-12 14:23:45 -04:00
Valentin Ruano-Rubio	ce8e93fa51	Made the AF prior probability distribution dynamic respect to the total-ploidy (added ploidy accross samples). Changes: -------- * Instead of calculate a fixed log10 prior array with a fix total likelihood we use a new component, the AFPriorProvider to generate the priors for different total plodies on demand; these are cached however so there is no unecessary recompute involved.	2014-09-12 14:23:37 -04:00
Valentin Ruano-Rubio	31e58ae4ec	Refactored AFCalc to remove unecessary capability limits allowing to deal with mixed ploidies and max-alt-allele number changes dynamically. Changes: -------- * Moved the AFCalcFactory.Calculation enum in a top level class AFCalculatorImplementation. * Given more reponsabilities to the enum like resolving the constructor method once per implementation and the best-model selection algorithm. * Removed test-code only fields and methods from AFCalc; just used to perform unit-testing and not any actual functionality of this component. * Removed the fixed ploidy constraint of GeneralPloidyExactAFCalc implementation... now can deal with mixed ploidies that may change per site and sample. * Removed the fixed maxAltAllele restriction by allowing resizing of the stateTracker structures. * Due to previous two points now call the the AFCalc object are passed the default-ploidy to assume in case some genotype in the input VC does not have it and the max-alt-allele. * Also due to those changes, removed the now totally useless 3 int parameters from all AFCalc constructors. * Cleaned the code a bit from no further used components and methods.	2014-09-12 14:17:36 -04:00
Eric Banks	7f6e526a87	Merge pull request #732 from broadinstitute/rp_ignore_all_filters_in_VQSR Added ignore all filters options to VQSR walkers	2014-09-12 00:17:20 -04:00
Ryan Poplin	48252897b4	Added ignore all filters options to VQSR walkers	2014-09-11 15:11:41 -04:00
Eric Banks	31cea25c36	Merge pull request #730 from broadinstitute/eb_inbreeding_coeff_unit_test Cleaned up and fleshed out unit tests for the Inbreeding Coefficient annotation class	2014-09-10 09:32:49 -04:00
Eric Banks	5e490362ca	Cleaned up and fleshed out unit tests for the Inbreeding Coefficient annotation class.	2014-09-08 11:40:39 -04:00
Eric Banks	78a61dc3e8	Merge pull request #729 from broadinstitute/eb_improve_dangling_head_merging_PT74221002 Improve the accuracy of dangling head merging in the HC assembler.	2014-09-08 10:28:51 -04:00
Eric Banks	cc175bad40	Improve the accuracy of dangling head merging in the HC assembler. Dangling head merging (like with tails) in now enabled by default. The --recoverDanglingHeads argument is now deprecated so that users know not to use it anymore. We now also allow the user to set the minimum branch length for merging. This will be different for exomes and RNA (see below). The other changes in the code itself: 1. We no longer allow an arbitrarily large number of mismatches in the dangling head for merging 2. The max number of mismatches allowed in a dangling head is proportional to the kmer size There will be a difference in the RNA calling pipeline. Instead of invoking '--recoverDanglingHeads' the user will instead want to use '--minDanglingBranchLength 0'. Below are the knowledgebase results of the master branch vs. this one. For NA12878 DNA Exome: master SNPS TRUE_POSITIVE 36722 master SNPS CALLED_NOT_IN_DB_AT_ALL 2699 master SNPS REASONABLE_FILTERS_WOULD_FILTER_FP_SITE 292 master SNPS FALSE_POSITIVE_SITE_IS_FP 70 branch SNPS TRUE_POSITIVE 36867 branch SNPS CALLED_NOT_IN_DB_AT_ALL 2952 branch SNPS REASONABLE_FILTERS_WOULD_FILTER_FP_SITE 387 branch SNPS FALSE_POSITIVE_SITE_IS_FP 94 As I discussed with Ryan in person, there are a good number of FPs that are called in the new code, but they nearly all have bad strand bias and should be easily filtered by VQSR. Note that there is no change for indels. For NA12878 RNA from Ami: master SNPS TRUE_POSITIVE 11055 master SNPS CALLED_NOT_IN_DB_AT_ALL 831 master SNPS REASONABLE_FILTERS_WOULD_FILTER_FP_SITE 44 master SNPS FALSE_POSITIVE_SITE_IS_FP 96 branch SNPS TRUE_POSITIVE 11113 branch SNPS CALLED_NOT_IN_DB_AT_ALL 874 branch SNPS REASONABLE_FILTERS_WOULD_FILTER_FP_SITE 47 branch SNPS FALSE_POSITIVE_SITE_IS_FP 92 Again, there's basically no change for indels.	2014-09-07 08:55:59 -04:00
Eric Banks	56a2554bb0	Merge pull request #720 from broadinstitute/pd_standardize_args Moved arguments controlling options in output files into the engine	2014-09-06 20:49:58 -04:00
Phillip Dexheimer	a35f5b8685	Moved arguments controlling options in output files into the engine * Arguments involved are --no_cmdline_in_header, --sites_only, and --bcf for VCF files and --bam_compression, --simplifyBAM, --disable_bam_indexing, and --generate_md5 for BAM files * PT 52740563 * Removed ReadUtils.createSAMFileWriterWithCompression(), replaced with ReadUtils.createSAMFileWriter(), which applies all appropriate engine-level arguments * Replaced hard-coded field names in ArgumentDefinitionField (Queue extension generator) with a Reflections-based lookup that will fail noisily during extension generation if there's an error	2014-09-05 21:18:11 -04:00
droazen	5c4a3eb89c	Merge pull request #727 from broadinstitute/ks_gatk_queue_package_test_updates Various fixes for package tests.	2014-09-05 10:17:32 -04:00
Eric Banks	0c336e0eb8	Merge pull request #728 from broadinstitute/rp_SOR_standard_annotation StrandOddsRatio is now a standard annotation.	2014-09-05 10:06:52 -04:00
Ryan Poplin	a45acdfb89	StrandOddsRatio is now a standard annotation.	2014-09-05 08:33:37 -04:00
Khalid Shakir	376592f423	Various fixes for package tests. Explicitly including gatk/queue test-jar artifacts in package test classpaths. SelectVariantsIntegrationTest#testInvalidJexl now resets the JexlEngine silent flag that VariantFiltration.initialize() toggles. External example no longer tries to unpack nonexistent gatk artifact jars during package tests.	2014-09-04 15:30:31 -04:00
rpoplin	9705df2f7a	Merge pull request #726 from broadinstitute/rp_haplotypeCaller_typo_fix fixing a few small typos in the HaplotypeCaller and related classes	2014-09-04 15:04:41 -04:00
Ryan Poplin	1b809268d5	fixing a few small typos in the HaplotypeCaller and related classes	2014-09-04 14:48:27 -04:00
kshakir	f53ea3b456	Merge pull request #725 from broadinstitute/ks_remove_ipflibrary_symlink Fixed old symlink reference in IPFLibraryQueueTest.	2014-09-04 22:14:02 +08:00
Khalid Shakir	f7c37eff06	Fixed old symlink reference in IPFLibraryQueueTest.	2014-09-04 22:13:17 +08:00
droazen	5c087a6e1f	Merge pull request #724 from broadinstitute/ks_remove_test_qscript_symbolic_links Removed symlink creation for tests and qscripts	2014-09-04 09:10:54 -04:00
Eric Banks	538537dbf1	Merge pull request #718 from broadinstitute/mf_rbp_fix Fix MNP merging code to work with explicit HP phase representation	2014-09-02 20:39:22 -04:00
Eric Banks	01e725cd1a	Merge pull request #723 from broadinstitute/eb_fix_rna_splitting_PT77878554 Make sure that the OverhangFixingManager (used for splitting RNA reads) ...	2014-09-02 20:39:01 -04:00
Menachem Fromer	10f9001738	Fix MNP merging code to work with explicit HP phase representation	2014-09-02 17:25:08 -04:00
Eric Banks	ff91ab8ba2	Make sure that the OverhangFixingManager (used for splitting RNA reads) handles unmapped reads.	2014-09-02 16:56:17 -04:00
Valentin Ruano Rubio	c7925f6e5c	Merge pull request #719 from broadinstitute/vrr_generalize_ploidy_in_genotype_gvcfs Adds support for omniploidy to GenotypeGVCFs and CombineGVCFs.	2014-09-02 16:51:02 -04:00
Valentin Ruano-Rubio	d363725b4b	Adds support for omniploidy to GenotypeGVCFs and CombineGVCFs. Same changes fixed the problem for GenotypeGVCFs and CombineGVCFs. Stories: - https://www.pivotaltracker.com/story/show/77626044 - https://www.pivotaltracker.com/story/show/77626854 Changes: - Generalized the code for the merging in GATKVariantContextUtils to cope with ploidy != 2. - GenotypeGVCFs now check that the input's ploidy conform to the '-ploidy' argument. - Moved out Refernce Confidence VC merging code from GATKVariantContextUtils so that we can keep new code in protected. Caveats: - GenotypeGVCFs only can deal with input files that have the same ploidy in all positions; the one that the user MUST indicate in the -ploidy argument (if different to the default 2). - CombineGVCFs won't necessarely complain if its passed mixed ploidy inputs but you won't be able to genotype it with GenotypeGVCFs. Test: - Removed deprecated unit tests for GATKVariantContextUtils. - Moved unit-tests regarding GVCF merging from GATKVariantContextUtilsUnitTest to ReferenceConfidenceVariantContextUtilsUnitTest. - Added unit test for new code for mapping genotype indices between allele index encoding in GenotypeLikelihoodCalculator. - GenotypeGVCFs and CombineGVCFs original integration test are unaffected by the change. - Added tetraploid run integration tests to check on non-diploid execution of GenotypeGVCFs and CombineGVCFs.	2014-09-02 15:06:47 -04:00
Eric Banks	fe86dafc41	Merge pull request #705 from broadinstitute/gg_simplify_gatkdocs_templates Changed the GATKDocs format to PHP	2014-09-02 06:28:26 -04:00
Khalid Shakir	fcb0eca203	Now passing in the path to the GATK directory to tests. Changed tests and scripts to use gatkdir full path instead of relative testdata/qscripts symbolic links. Although symlinks not created, left the symlink deletion script execution with a comment about future removal. Re-enabled example UG pipeline queue test. Replaced all hardcoded strings of {public,private}/testdata with BaseTest variables. Refactored temp list creation method from ListFileUtilsUnitTest to BaseTest.createTempListFile. Removed list files with hardcoded paths, now using createTempListFile instead with private test dir variable.	2014-09-02 01:40:59 +08:00
kshakir	9477a6ab1a	Merge pull request #722 from broadinstitute/ks_mlinderm_taggable_RodBindingCollection Update extension generator to recognize RodBindingCollection as 'taggable'	2014-09-01 18:07:07 +08:00
Michael Linderman	380cd67146	Update extension generator to recognize RodBindingCollection as 'taggable' Signed-off-by: Khalid Shakir <kshakir@broadinstitute.org>	2014-09-01 12:19:44 +08:00
Eric Banks	46b0c18603	Merge pull request #721 from broadinstitute/ks_save_analyze_covariates_integration_test_inputs Don't delete "after" files in AnalyzeCovariatesIntegrationTest	2014-08-29 18:44:40 -04:00
Khalid Shakir	2d28972c88	The 'after' files are @Input files and commited in git, so don't delete them after tests.	2014-08-30 03:04:54 +08:00
Eric Banks	b654590ed6	Merge pull request #717 from broadinstitute/eb_change_phasing_of_hom_vars Changed the functionality of the physical phasing in the HC: now hom var...	2014-08-25 21:41:11 -04:00
Eric Banks	5b087c9897	Changed the functionality of the physical phasing in the HC: now hom vars are output as 0\|1. We do this for technical reasons, mostly because we don't genotype in the HC anymore; it's all done downstream by GenotypeGVCFs so we can't be sure that the genotype will be hom var. Also, there are steps in the downstream pipeline where genotypes can change, so assuming anything in the HC is a bad idea, and if we have phasing info in the het state, we want to propagate that forward. Now, PGT tag fixing happens downstream in GenotypeGVCFs. While I was in there I also cleaned up the code a bit and fixed a bug where annotation was happening before genotype creation when using the --includeNonVariantSites argument. Added tests accordingly.	2014-08-25 21:40:14 -04:00
Valentin Ruano Rubio	9324121af4	Merge pull request #716 from broadinstitute/vrr_broken_md5s Fixes some missmerged md5 updates from a previous merge into master	2014-08-25 10:30:00 -04:00
Valentin Ruano-Rubio	6dc5cf0be0	Fixes some missmerged md5 updates from a previous merge into master	2014-08-24 20:47:07 -04:00
Eric Banks	9009c1e996	Merge pull request #715 from broadinstitute/vrr_disable_physical_phasing_for_nondiploid_hc Disable physical phasing for non-diploid HC calling.	2014-08-23 20:58:51 -04:00
Eric Banks	34e5cce553	Merge pull request #714 from broadinstitute/pd_hc_samplename Add the --sample_name argument to HaplotypeCaller	2014-08-23 20:57:44 -04:00
Valentin Ruano-Rubio	6695aeafd9	Disable physical phasing for non-diploid HC calling. Story: https://www.pivotaltracker.com/story/show/77452256 Changes: If ploidy != 2, disable physical phasing and log an info message to let the user know. Tests: Change md5s affected by this change.	2014-08-23 10:52:07 -04:00
Phillip Dexheimer	931890915f	Add the --sample_name argument to HaplotypeCaller * This is a shortcut for people who have multi-sample BAMs but would like to use GVCF mode. Rather than creating single-sample BAMs with PrintReads, one could use the --sample_name argument to HaplotypeCaller to specify the single sample to make calls on * Completes PT 73075482	2014-08-22 23:22:03 -04:00
Valentin Ruano Rubio	2b129ba707	Merge pull request #713 from broadinstitute/vrr_mlpsacaf_standalone_annotations Created the stand-alone AC and AF annotation AlleleCountsBySample	2014-08-22 22:19:28 -04:00
Valentin Ruano-Rubio	fc5ce4b662	Created the stand-alone AC and AF annotation AlleleCountBySample Story: https://www.pivotaltracker.com/story/show/77250524 Changes: - Remove the annotating code in GeneralPloidyExactAFCalc (GPEAFC) class. - Added the asAlleleList to GenotypeAlleleCounts class and get (GPEAFC) to use that instead of implementing its own (nicer and more reusable code). - Removed the explicit addition of AlleleCountBySample fields to the VCF header by the walker initialize - Added utility methods in Utils to wrap and int[] array into a List<Integer>, and double[] array into a List<Double> efficiently. Test: - Added unit-testing for asAlleleList in GenotypeAlleleCountsUnitTest (within testFirst and testNext). - Added unit-testing for new methods in Utils : asList(int[]) and asList(double[]) - Changed UG General Ploidy test to add explicitly those annotations. - Non-trivial changes in integration tests involving non-diploid runs (namelly haploid and tetraploid) as they are not showing those annotations anylonger, so the MD5s have been changed accordingly.	2014-08-22 20:33:25 -04:00
Eric Banks	36bdfa3918	Merge pull request #712 from broadinstitute/eb_physical_phasing_bug_PT77248992 Fixing bug in the physical phasing code, found by Valentin.	2014-08-21 15:25:51 -04:00
Eric Banks	b1cb6196be	Fixing bug in the physical phasing code, found by Valentin. It turns out that there can be some really complex situations even with a single sample where there are lots of unphasable hets around a hom. Previously we were trying to phase each of the hets against the hom, but that wasn't correct. Instead we now detect that situation and don't attempt to phase anything. Added a unit test to cover this situation.	2014-08-21 15:24:09 -04:00
rpoplin	e60dd77362	Merge pull request #711 from broadinstitute/ldg_deNovoAnnotation Add bells and whistles for Genotype Refinement Pipeline	2014-08-21 15:06:42 -04:00
Laura Gauthier	9a5da41dd4	Add bells and whistles for Genotype Refinement Pipeline New annotation for low= and high-confidence de novos (only annotates biallelics) FamilyLikelihoodsUtils now add joint likelihood and joint posterior annotations Restrict population priors based on discovered allele count to be valid for 10 or more samples.	2014-08-21 11:20:40 -04:00
Valentin Ruano Rubio	0c25eb7163	Merge pull request #709 from broadinstitute/vrr_fix_general_diploid_exact_af_index_out_of_bounds_bug Fix for the GeneralPloidyExactAFCalc implementation that was preventing -ploidy != 2 GVCF/BP_RESOLUTION output to work. Story: https://www.pivotaltracker.com/story/show/74471252 Tests: Enabled GVCF tests with ploidy != 2 and other checking for the original ArrayIndexOutOfBounds exception.	2014-08-20 16:57:51 -04:00
Valentin Ruano-Rubio	d31c5536aa	Fixed the bug first by indicating the actual possible number of alternatives alleles considering the extra <NON_REF> and second by resizing the StateTracker capacity when invoked by GeneralPloidyExactAFCalc deep within its implementation of computeLog10PNonRef which is ultimatelly what get rids of the exception. Story: https://www.pivotaltracker.com/story/show/74471252	2014-08-20 14:42:42 -04:00

1 2 3 4 5 ...

13639 Commits (95b45443ae584cbc01250a3359b3bdc1906f8dea) All Branches Search

13639 Commits (95b45443ae584cbc01250a3359b3bdc1906f8dea)

All Branches