gatk-3.8

Commit Graph

Author	SHA1	Message	Date
Eric Banks	623b36fbc4	Add header lines for AC,AF, and AN tags	2012-05-02 15:33:34 -04:00
Guillermo del Angel	429800a192	Fix corner case rounding issue in MathUtils unit test: 10^logFactorial(4)) was 23.999999... which if cast directly yielded 23 - so, do pre-rounding to ensure correct integer result if caller will cast value.	2012-05-02 09:57:06 -04:00
Guillermo del Angel	76a95fdedf	Full implementation of multiallelic exact model for pools. Still super-linear so not useable at scale but it should be a gold standard to compare to. Unit tests are not exhaustive yet, will be expanded to provide better test coverage. Small inconsequential optimization in MathUtils: we're already caching log10(factorial(n)) for large n, so might as well use the cached values to compute binomial and multinomial coefficients instead of the log-gamma approximation which is more expensive (doesn't seem to save much time either in PoolCaller nor in UG though).	2012-05-02 09:24:28 -04:00
Joel Thibault	4d732fa586	Move all MongoDB files into private/java/src/org/broadinstitute/sting/mongodb	2012-05-01 18:23:51 -04:00
Eric Banks	619a69a5f1	As promised in the release notes for 1.6, I am removing the old deprecated genotyping framework revolving around the misordering of alleles and have moved the fixed version in its place in preparation for release 1.7 (or 2.0?).	2012-05-01 16:18:24 -04:00
Joel Thibault	c255dd5917	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-05-01 16:10:38 -04:00
Ryan Poplin	51af61b5d7	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-05-01 16:07:23 -04:00
Ryan Poplin	fc55dcec3c	Unfortunately the reverse trimming of alleles still doesn't work with mixed records in some corner cases. Turning it off for now.	2012-05-01 16:02:36 -04:00
Ryan Poplin	20a0078f23	Merging active regions across shard boundries if they are contiguous, have the same active status and don't grow too big.	2012-05-01 15:51:36 -04:00
Eric Banks	0f3af9555b	Adding an option to SelectVariants which allows the user to re-genotype through the exact model (if PLs are present) the samples in order to recalculate the QUAL and genotypes. This is really the correct way to select a subset of samples, especially when originally called from low coverage data. Also added integration test to cover this case.	2012-05-01 14:58:06 -04:00
Joel Thibault	aa4d41cce0	Minor cleanup before push	2012-05-01 14:16:44 -04:00
Joel Thibault	b101b9c30b	Add Mongo switch	2012-05-01 14:00:48 -04:00
Joel Thibault	1b609e9075	Move Mongo to server couchdb	2012-05-01 13:59:47 -04:00
Joel Thibault	fd57d27f45	Move MongoDB connection handling to a separate class	2012-05-01 13:59:37 -04:00
Joel Thibault	db3cd1abd5	Use 2 MongoDB collections (tables): one for INFO/attributes, one for samples/genotypes.	2012-05-01 13:57:23 -04:00
Joel Thibault	04e1be9106	Better handling of Mongo errors + exceptions	2012-05-01 13:57:23 -04:00
Joel Thibault	ca737479cf	Query for stop locations because we don't have that information in the reference	2012-05-01 13:57:23 -04:00
Joel Thibault	1cda87a4ad	Set ROD priority list to input	2012-05-01 13:57:23 -04:00
Joel Thibault	a7fe847faf	Set the priority list and don't bother combining if not needed	2012-05-01 13:57:23 -04:00
Joel Thibault	f739305f43	Combine the variants found at a location	2012-05-01 13:57:23 -04:00
Joel Thibault	020f884d5a	Use new key of source ROD plus alleles	2012-05-01 13:57:23 -04:00
Joel Thibault	221ce9c3d6	Add alleles to the primary key	2012-05-01 13:57:23 -04:00
Joel Thibault	3198ce5471	Can have multiple variants at a location	2012-05-01 13:57:22 -04:00
Joel Thibault	11ed8e61c9	Add referenceBaseForIndel to the Mongo VariantContext objects	2012-05-01 13:53:44 -04:00
Joel Thibault	7ed0ee7ed0	Skip locations with no genotypes instead of throwing a NPE	2012-05-01 13:53:44 -04:00
Joel Thibault	4bdfeacdaa	Handle multiple samples/genotypes per location TODO: sample selection	2012-05-01 13:53:43 -04:00
Joel Thibault	1f7c628796	Insert the ROD filename into MongoDB as part of the primary key	2012-05-01 13:53:43 -04:00
Joel Thibault	bb8a6e9b0a	Initial test of write and read from MongoDB	2012-05-01 13:53:43 -04:00
David Roazen	c0084c741b	Pilot BCF2 Implementation: Checkpointing the code * Not working yet, still very much a work-in-progress with lots of placeholders * Needed to check this in to enable possible collaboration, since it's going slower than anticipated and the conference deadline looms.	2012-05-01 12:23:10 -04:00
Christopher Hartl	7d029b9a28	Merge branch 'master' of ssh://ni.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-30 12:16:30 -04:00
Christopher Hartl	944a7d815e	Bringing VQSRV3 up to date. Lots of new features (un-classifying the worst-performing training sites, treating the x% best/worst sites as postive/negative points, ability to pass in a monomorphic track to see ROC curves output). Minor changes to AlleleBalance: weighted average was incorrectly specified (using logscale actually biased the average towards the AB of low-quality genotypes), and breaking out AB by het, hom, and diploid to bring it in line with some (private) changes to the indel likelihood model that (correctly) computes these values for indels.	2012-04-28 11:31:03 -04:00
Ryan Poplin	54a9bc2da2	Bug fix in reverse trim alleles for the case of mixed records that become non-mixed after subsetting the alleles.	2012-04-28 09:12:26 -04:00
Ryan Poplin	e332aeaf70	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-27 16:21:21 -04:00
Ryan Poplin	2b5dd28550	Bug fix in reverse trim alleles for the case of mixed records.	2012-04-27 16:21:02 -04:00
Mauricio Carneiro	1db2d1ba82	Do not add the first and last 4 cycles to the recalibration tables.	2012-04-27 15:18:07 -04:00
Mauricio Carneiro	08dbd756f3	Quick QC walkers to look at the error profile of indels in the read	2012-04-27 15:18:07 -04:00
Guillermo del Angel	730208133b	Several fixes and improvements to Pool caller with ancillary test functions (not done yet): a) Utility class called Probability Vector that holds a log-probability vector and has the ability to clip ends that deviate largely from max value. b) Used this class to hold site error model, since likelihoods of error model away from peak are so far down that it's not worth computing with them and just wastes time. c) Expand unit tests and add an exhaustive test for ErrorModel class. d) Corrected major math bug in ErrorModel uncovered by exhaustive test: log(e^x) is NOT x if log's base = 10. e) Refactored utility functions that created artificial pileups for testing into separate class ArtificialPileupTestProvider. Right now functionality is limited (one artificial contig of 10 bp), can only specify pileups in one position with a given number of matches and mismatches to ref) but functionality will be expanded in future to cover more test cases. f) Use this utility class for IndelGenotypeLikelihoods unit test and for PoolGenotypeLikelihoods unit test (the latter testing functionality still not done). g) Linearized implementation of biallelic exact model (very simple approach, similar to diploid exact model, just abort if we're past the max value of AC distribution and below a threshold). Still need to add unit tests for this and to expand to multiallelic model. h) Update integration test md5's due to minor differences stemming from linearized exact model and better error model math	2012-04-27 14:41:17 -04:00
Eric Banks	0439047269	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-27 10:49:45 -04:00
Eric Banks	05b44dd017	The genotypeCounts array wasn't always being initialized before it was accessed, leading to a NPE (which got caught and thrown as a JEXL expression when used in selection). Added unit test to cover all genotype count methods.	2012-04-27 10:49:36 -04:00
Khalid Shakir	9801dd114f	Bug fix for: https://getsatisfaction.com/gsa/topics/problem_with_indelrealigner_and_l_unmapped The GATK -L unmapped is for GenomeLocs with SAMRecord.NO_ALIGNMENT_REFERENCE_NAME, not SAMRecord.getReadUnmappedFlag() Previously unmapped flag reads in the last bin were being printed while also seeking for the reads without a reference contig.	2012-04-27 09:58:38 -04:00
Guillermo del Angel	972d6531b6	Corner case fix for indel GL computation: sometimes (depending on surrounding context) reads which are not informative of two candidate haplotypes end up having marginally higher likelihoods with one haplotype as opposed to another, depending on uncertainty on alignments in surrounding regions. So, a sample whose GL is -0.0001,-0.0005,-0.001 may have its genotype set to 1/1 due to this statistical noise. We already have a tolerance comparing max(gl)-min(gl) to avoid genotyping, so this tolerance is now increased from 0.001 to 0.1 (equivalent to 1 PL unit) to avoid genotyping a sample if all PLs are within this threshold. Changed 2 integration test md5s that hit this case.	2012-04-26 10:15:26 -04:00
Laurent Francioli	219b0a128b	PED support for ChromosomeCounts annotation Signed-off-by: Eric Banks <ebanks@broadinstitute.org>	2012-04-25 12:50:04 -04:00
Laurent Francioli	19d5213d5a	Added function to get founders IDs in SampleDB Signed-off-by: Eric Banks <ebanks@broadinstitute.org>	2012-04-25 12:49:36 -04:00
Mauricio Carneiro	902277856e	fix for RBP getPileupsForSamples() do not differentiate per sample pileups from generic pileups. Do the same for both -- it's O(n) either way.	2012-04-24 17:20:30 -04:00
Mauricio Carneiro	82b4798913	CountBasesWalker -- a quick QC walker.	2012-04-24 17:20:30 -04:00
Mauricio Carneiro	e440d0ce69	BQSR triage #4 * fixed queue script plot file names * updated the ReadGroupCovariate to use the platform unit instead of sample + lane. * fixed plotting of marginalized reported qualities	2012-04-24 17:19:54 -04:00
Eric Banks	d6277b70d8	Forgot to consider the optimized case in hasAllele	2012-04-24 11:32:28 -04:00
Eric Banks	91bad244d5	Using a VCF whose ALT is the reference in GGA mode is a User Error	2012-04-24 11:08:37 -04:00
Eric Banks	74ad008163	Adding VariantContext.hasAlternateAllele functionality	2012-04-24 11:07:46 -04:00
Eric Banks	66f3315548	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-24 09:39:55 -04:00
Eric Banks	bcb93dda5f	Fixing docs (rank sum test values are not phred-scaled)	2012-04-24 09:39:42 -04:00
Mauricio Carneiro	e39a59594a	BQSR triage and test routines * updated BQSR queue script for faster turnaround * implemented plot generation for scatter/gatherered runs * adjusted output file names to be cooperative with the queue script * added the recalibration report file to the argument table in the report * added ReadCovariates unit test -- guarantees that all the covariates are being generated for every base in the read * added RecalibrationReport unit test -- guarantees the integrity of the delta tables	2012-04-23 11:23:00 -04:00
Eric Banks	a733723439	Merged bug fix from Stable into Unstable	2012-04-23 10:30:30 -04:00
Eric Banks	2761da975e	Handle null VCs (which can arise when indels are present in the file)	2012-04-23 10:30:00 -04:00
Eric Banks	63aa79df82	Slightly better error message	2012-04-23 09:37:28 -04:00
Eric Banks	7b5fbf9567	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-23 09:34:08 -04:00
Eric Banks	4edb005411	Catch poorly formatted PL/GL fields	2012-04-23 09:33:50 -04:00
Ryan Poplin	35bb55f562	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-22 13:23:36 -04:00
Ryan Poplin	18e4532d10	Turning down the amount of assembly graph pruning slightly in the case of low coverage.	2012-04-22 13:23:24 -04:00
Eric Banks	1f23d99dfa	If we are subsetting alleles in the UG (either because there were too many or because some were not polymorphic), then we may need to trim the alleles (because the original VariantContext may have had to pad at the end). Thanks to Ryan for reporting this. Only one of the integration tests had even partially covered this case, so I added one that did.	2012-04-20 17:00:05 -04:00
Eric Banks	4b81c75642	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-20 14:30:19 -04:00
Eric Banks	f1c5510ec0	When running SelectVariants with the excludeNonVariants option, remove alleles from the ALT field that are no longer polymorphic.	2012-04-20 14:30:04 -04:00
Ryan Poplin	a1596791af	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-20 14:03:04 -04:00
Ryan Poplin	a57295eb75	Fixing a bug when breaking up active regions where the resulting regions would overlap by one base. Adding quality score manipulation from the UG into the haplotype caller (qual capped by mapping quality, min qual threshold).	2012-04-20 14:02:55 -04:00
Guillermo del Angel	de68363c23	Removed experimental feature (aka hack) that was meant for 1000G consensus but remained in VQSR data manager - QD was being scaled by indel length. There's no evidence any more that QD is length-dependent, neither in CEU trio data nor in latest 1000G P2 calls	2012-04-20 10:58:34 -04:00
Guillermo del Angel	d2488dfb81	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-19 19:40:03 -04:00
Guillermo del Angel	c44c7b9a97	Restored optimization in Pair HMM only to compute HMM matrices starting in index where haplotypes start to diverge - saves about 15-20% of runtime which is what we lost by disabling banding in latest version, so runtime should be now about the same as what it was before refactoring. Output is bit-true to previous commit	2012-04-19 19:39:43 -04:00
Mauricio Carneiro	0f8c77391d	BQSR bug triage #3 * fixed context covariate famous "off by one" error * reduced maximum quality score to Q50 (following Eric/Ryan's suggestion) * remove context downsampling in BQSR R script	2012-04-19 17:31:04 -04:00
Khalid Shakir	df5dd841af	AC strat now checks if evals will be merged before throwing an error on multiple eval files. Minor tweaks to WGP script based on new recal VCF format.	2012-04-19 16:08:55 -04:00
Guillermo del Angel	1ae2ab5b63	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-19 12:50:29 -04:00
Guillermo del Angel	0e6e0cb907	Merging bug fixes	2012-04-19 12:49:30 -04:00
Eric Banks	79272c5e15	Thanks to Menachem for pointing out that the docs for genotyping_mode and output_mode were the same (and unclear). Fixed.	2012-04-19 12:48:09 -04:00
Guillermo del Angel	02ff930f6a	My changes	2012-04-19 12:45:18 -04:00
Eric Banks	2485cef5b8	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-19 11:46:06 -04:00
Eric Banks	76a6e37f4f	Don't output callability metrics by default anymore; one can still have them output to the 'metrics' file (which is now @Hidden because they are really for GSA use). Added a TODO to move UG from @By reference to reads and rods once LIBS is cleaned up.	2012-04-19 11:45:56 -04:00
Ryan Poplin	1ea4e48a27	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-19 11:32:32 -04:00
Ryan Poplin	11001ab9a2	Adding option to HaplotypeCaller to genotype the events on the chosen haplotypes as independent events. The filtered reads are now kept around so they can be passed to the variant annotations. Unfortunately the filtered reads aren't assigned a likelihood yet so they are all thrown in the Allele.NO_CALL bin.	2012-04-19 11:32:10 -04:00
Mauricio Carneiro	eb22cd7222	Unit test to guarantee BQSR sequential calculation accuracy This test brings together the old and the new BQSR, building a recalibration table using the two separate frameworks and performing the recalibration calculation using the two different frameworks for 10,000+ bases and asserting that the calculations match in every case.	2012-04-19 09:33:40 -04:00
Mauricio Carneiro	68d0211fa1	Improved BQSR plotting and some new parameters * Refactored CycleCovariate to be a fragment covariate instead of a per read covariate * Refactored the CycleCovariateUnitTest to test the pairing information * Updated BQSR Integration tests accordingly * Made quantization levels parameter not hidden anymore * Added hidden option to keep intermediate plotting files for debug purposes (they're automatically deleted) * Added hidden option not to generate the plots automatically (important for scatter/gathering)	2012-04-19 09:31:41 -04:00
Guillermo del Angel	143e92b797	Rebasing	2012-04-18 20:05:43 -04:00
Guillermo del Angel	82efd4457e	Revert some bad merge changes	2012-04-18 16:35:09 -04:00
Guillermo del Angel	31c394d588	Resolve merge conflicts	2012-04-18 16:25:03 -04:00
Ryan Poplin	4999ae87ad	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-18 15:02:42 -04:00
Ryan Poplin	dcc4871468	minor misc optimizations to PairHMM	2012-04-18 15:02:26 -04:00
Eric Banks	d3c84e7b1f	This should be a User Error since it's provided from the DoC command-line arguments	2012-04-18 13:09:23 -04:00
Eric Banks	392f1903f7	Handling some of the NumberFormatExceptions seen via Tableau that are really user errors.	2012-04-18 12:57:37 -04:00
Ryan Poplin	8a84456626	Following Eric's awesome update to change the VQSR recal file into a VCF file, the ApplyRecalibration step is now scatter/gather-able and tree reducible.	2012-04-18 11:24:04 -04:00
Eric Banks	4448a3ea76	Final tweaks. Added an integration test to cover the case of SNPs and indels that start at the same position.	2012-04-17 23:54:10 -04:00
Eric Banks	c1f52b773a	Minor tweaks and updated integration tests MD5s	2012-04-17 23:17:28 -04:00
Eric Banks	6d03bce0d3	Important refactoring of the VQSR recal file format: we now use a VCF instead of a CSV file. The most important reason for this change is that we no longer need to read the entire recal file into memory up front in ApplyRecalibration. For 1000G calling this was prohibitive in terms of memory requirements. Now we go through the rod system and pull in just the records we need at a given position. As an added bonus, once BCF2 is live we can drastically cut down the sizes of these recal files (which can grow large for whole genome calling).	2012-04-17 22:38:18 -04:00
Mauricio Carneiro	46a212d8e9	Added "simplify reads" option to PrintReads.	2012-04-17 19:32:34 -04:00
Mauricio Carneiro	f0c81b59b0	Implementation of the new BQSR plotting infrastructure * removed low quality bases from the recalibration report. * refactored the Datum (Recal and Accuracy) class structure * created a new plotting csv table for optimized performance with the R script * added a datum object that carries the accuracy information (AccuracyDatum) for plotting * added mean reported quality score to all covariates * added QualityScore as a covariate for plotting purposes * added unit test to the key manager to operate with one required covariate and multiple optional covariates * integrated the plotting into BQSR (automatically generates the pdf with the recalibration tearsheet)	2012-04-17 19:23:55 -04:00
Ryan Poplin	952280bef1	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-17 17:00:14 -04:00
Ryan Poplin	cf705f6c62	Adding read position rank sum test to the list of annotations that get produced with the HaplotypeCaller	2012-04-17 17:00:00 -04:00
Eric Banks	13c800417e	Handle NPE in UG indel code: deletions immediately preceding insertions were not handled well in the code.	2012-04-17 15:51:23 -04:00
Guillermo del Angel	c78b0eee3a	Refactoring/fixing up UG HMM code: a) Make code use PairHMM class instead of having duplicated code. That way UG and HaplotypeCaller now use same core code. Changes to be able to do this: 1. Compute context-dependent GOP as a function of read, not of haplotype, b) Extracted code to initialize HMM arrays into separate method, c) Move PairHMM class and unit test to public, d) Reenable banded code in PairHMM, inverted sense of flag (true=enable feature) but leave off in HaplotypeCaller.	2012-04-17 14:22:48 -04:00
Khalid Shakir	91cb654791	AggregateMetrics: - By porting from jython to java now accessible to Queue via automatic extension generation. - Better handling for problematic sample names by using PicardAggregationUtils. GATKReportTable looks up keys using arrays instead of dot-separated strings, which is useful when a sample has a period in the name. CombineVariants has option to suppress the header with the command line, which is now invoked during VCF gathering. Added SelectHeaders walker for filtering headers for dbGAP submission. Generated command line for read filters now correctly prefixes the argument name as --read_filter instead of -read_filter. Latest WholeGenomePipeline. Other minor cleanup to utility methods.	2012-04-17 11:45:32 -04:00
Ryan Poplin	1a2e92f8db	Merged bug fix from Stable into Unstable	2012-04-17 10:23:05 -04:00
Ryan Poplin	adad76b36f	Fixing NPE in VQSR for the case of very small callsets.	2012-04-17 10:20:43 -04:00
Mark DePristo	23ccf772d4	IndelSummary now emits all of the underlying counts for ratios, percentages, etc it computes	2012-04-13 17:00:36 -04:00
Mark DePristo	84d1e8713a	Infrastructure for combining VariantEvaluations -- Not hooked up yet, so the output of VariantEval should be the same as before -- Implemented a VariantEvalUnitTest that tests the low level strat / eval combinatorics and counting routines -- Better docs throughout	2012-04-13 17:00:36 -04:00
Mark DePristo	38986e4240	Documentation for StratificationManager	2012-04-13 17:00:36 -04:00
Mark DePristo	ab06d53867	Useful test constructor or Unit tests in RefMetaDataTracker	2012-04-13 17:00:36 -04:00
Mark DePristo	285e61a227	Bugfix for IndelSummary -- multi allelic count should be % not ratio	2012-04-13 17:00:35 -04:00
Mark DePristo	e6d5cb46d2	Improvements and bugfixes to IndelSummary -- Now properly includes both bi and multi-allelic variants. These are actually counted as well, and emitted as counts and % of sites with multiple alleles -- Bug fix for gold standard rate	2012-04-13 17:00:35 -04:00
Mark DePristo	bfa966a4e9	Bugfix for OneBPIndel -- Previously was only including 1 bp insertions in stratification	2012-04-13 17:00:35 -04:00
Mark DePristo	2aa2d9aec0	Merged bug fix from Stable into Unstable	2012-04-13 09:25:43 -04:00
Mark DePristo	27e7e17dc7	New way to handle exceptions in multi-threaded GATK -- HMS no longer tries to grab and throw all exceptions. Exceptions are just thrown directly now. -- Proper error handling is handled by functions in HMS, which are used by ShardTraverser and TreeReducer -- Better printing of stack traces in WalkerTest	2012-04-13 09:23:33 -04:00
Eric Banks	818e8c2fb9	Resolving merge conflicts	2012-04-12 15:19:44 -04:00
Eric Banks	0dd571928d	Let's not have the indel model emit more than the max possible number of genotypable alt alleles (since we may not be able to subset down to the best ones).	2012-04-12 15:16:29 -04:00
Eric Banks	f77a6d18b8	Bad conflict merge before	2012-04-12 09:56:49 -04:00
Eric Banks	33a8bdd75f	Resolving merge conflicts	2012-04-12 09:51:55 -04:00
Eric Banks	b659b16b31	Generate User Error for bad POS value	2012-04-12 09:49:35 -04:00
Eric Banks	cc71baf691	Don't allow users to try to genotype more than the max possible value (catch and throw a User Error at startup). Better docs explaining that users shouldn't play with this value unless they know what they are doing.	2012-04-12 09:18:44 -04:00
Eric Banks	5bf9dd2def	A framework to get annotations working in the HaplotypeCaller (and ART walkers in general). Adding support for active-region-based annotation for most standard annotations. I need to discuss with Ryan what to do about tests that require offsets into the reads (since I don't have access to the offsets) like e.g. the ReadPosRankSumTest. IMPORTANT NOTE: this is still very much a dev effort and can only be accessed through private walkers (i.e. the HaplotypeCaller). The interface is in flux and so we are making no attempt at all to make it clean or to merge this with the Locus-Traversal-based annotation system. When we are satisfied that it's working properly and have settled on the proper interface, we will clean it up then.	2012-04-11 16:22:12 -04:00
Guillermo del Angel	f9f8589692	Refactoring/fixing up UG HMM code: a) Make code use PairHMM class instead of having duplicated code. That way UG and HaplotypeCaller now use same core code. Changes to be able to do this: 1. Compute context-dependent GOP as a function of read, not of haplotype, b) Extracted code to initialize HMM arrays into separate method, c) Move PairHMM class and unit test to public, d) Reenable banded code in PairHMM, inverted sense of flag (true=enable feature) but leave off in HaplotypeCaller.	2012-04-11 13:56:51 -04:00
Eric Banks	7aa654d13f	New interface for some dev work that Ryan and I are doing; only accessible from private walkers right now	2012-04-11 13:49:09 -04:00
Eric Banks	dc90508104	Adding a new annotation to UG calls: NDA = number of discovered (but not necessarily genotyped) alleles for the site. This could help downstream analysis esp. of indels for wonky sites (since we only use the top 2-3 alleles). Not enabled by default but we can change that if this turns out to be useful.	2012-04-11 13:47:10 -04:00
Eric Banks	f560611fe8	Merged bug fix from Stable into Unstable	2012-04-10 22:26:53 -04:00
Eric Banks	f46f7d0590	Fix the stats coming out of FlagStat. I will add an integration test in unstable	2012-04-10 22:26:10 -04:00
Mauricio Carneiro	cd842b650e	Optimizing DiagnoseTargets * Fixed output format to get a valid vcf * Optimzed the per sample pileup routine O(n^2) => O(n) pileup for samples * Added support to overlapping intervals * Removed expand target functionality (for now) * Removed total depth (pointless metric)	2012-04-10 17:43:59 -04:00
Ryan Poplin	e3cc7cc59c	Resolving merge conflict.	2012-04-10 14:50:27 -04:00
Ryan Poplin	a4634624b7	There are now three triggering options in the HaplotypeCaller. The default (mismatches, insertions, deletions, high quality soft clips), an external alleles file (from the UG for example), or extended triggers which include low quality soft clips, bad mates and unmapped mates. Added better algorithm for band pass filtering an ActivityProfile and breaking them apart when they get too big. Greatly increased the specificity of the caller by battening down the hatches on things like base quality and mapping quality thresholds for both the assembler and the likelihood function.	2012-04-10 14:48:23 -04:00
Eric Banks	10e74a71eb	We now allow arbitrary annotations other than dbSNP (e.g. HM3) to come out of the Unified Genotyper. This was already set up in the Variant Annotator Engine and was just a matter of hooking UG up to it. Added integration test to ensure correct behavior.	2012-04-10 12:30:35 -04:00
Mark DePristo	b43d21056b	Merged bug fix from Stable into Unstable	2012-04-10 09:42:09 -04:00
Mark DePristo	6885e2d065	UserException fixes for GATK_logs recent errors -- SamFileReader.java:525 -- BlockCompressedInputStream:376 These were both instances were we weren't catching and rethrowing picard exceptions as UserExceptions.	2012-04-10 07:37:42 -04:00
Mark DePristo	8507cd7440	Throw UserException for bad dict / chain files	2012-04-10 07:22:43 -04:00
Ryan Poplin	cd9bf1bfc3	Changing IndelSummary eval module so that PostCallingQC.scala can run with MIXED-record VCFs.	2012-04-10 00:22:40 -04:00
Roger Zurawicki	9ece93ae9c	DiagnoseTargets now outputs a VCF file - refactored the statistics classes - concurrent callable statuses by sample are now available. Signed-off-by: Mauricio Carneiro <carneiro@broadinstitute.org>	2012-04-09 16:40:20 -04:00
Guillermo del Angel	719ec9144a	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-09 14:53:19 -04:00
Guillermo del Angel	550179a1f7	Major refactorings/optimizations of pool caller, output still bit-true to older version: a) Move DEFAULT_PLOIDY from UnifiedGenotyperEngine to VariantContextUtils. b) Optimize iteration through all possible allele combinations. c) Don't store log PL's in hashmap from allele conformations to double, it was too slow. Things can still be optimized much more down the line if needed. d) Remove remaining traces of genotype priors.	2012-04-09 14:53:05 -04:00
Eric Banks	ea4300d583	Refactoring so that Unified Argument Collection doesn't use deprecated classes.	2012-04-09 13:45:17 -04:00
Eric Banks	6ddf2170b6	More efficient implementation of the sum of the allele frequency posteriors matrix using a pre-allocated cache as discussed in group meeting last week. Now, when the cache is filled, we safely collapse down to a single value in real space and put the un-re-centered log10 value back into the front of the cache. Thanks to all for the help and advice.	2012-04-09 11:46:16 -04:00
Mauricio Carneiro	87e6bea6c1	Adding engine capability to quantize qualities. * Added parameter -qq to quantize qualities using a recalibration report * Added options to quantize using the recalibration report quantization levels, new nLevels and no quantization. * Updated BQSR scripts to make use of the new parameters	2012-04-08 21:07:51 -04:00
Mark DePristo	45fc0ea98d	Improvements to indel analysis capabilities of VariantEval -- Now calculates the number of Indels overlapping gold standard sites, as well as the percent of indels overlapping gold standard sites -- Removed insertion : deletion ratio for 1 bp event, replaced it with 1 + 2 : 3 bp ratio for insertions and deletions separately. This is based on an old email from Mark Daly: // - Since 1 & 2 bp insertions and 1 & 2 bp deletions are equally likely to cause a // downstream frameshift, if we make the simplifying assumptions that 3 bp ins // and 3bp del (adding/subtracting 1 AA in general) are roughly comparably // selected against, we should see a consistent 1+2 : 3 bp ratio for insertions // as for deletions, and certainly would expect consistency between in/dels that // multiple methods find and in/dels that are unique to one method (since deletions // are more common and the artifacts differ, it is probably worth looking at the totals, // overlaps and ratios for insertions and deletions separately in the methods // comparison and in this case don't even need to make the simplifying in = del functional assumption -- Added a new VEW argument to bind a gold standard track -- Added two new stratifications: OneBPIndel and TandemRepeat which do exactly what you imagine they do -- Deleted random unused functions in IndelUtils	2012-04-06 16:07:46 -04:00
Mark DePristo	52ef4a3e26	Function to compute whether a VariantContext indel is part of a TandemRepeat Returns true iff VC is an non-complex indel where every allele represents an expansion or contraction of a series of identical bases in the reference. The logic of this function is pretty simple. Take all of the non-null alleles in VC. For each insertion allele of n bases, check if that allele matches the next n reference bases. For each deletion allele of n bases, check if this matches the reference bases at n - 2 n, as it must necessarily match the first n bases. If this test returns true for all alleles you are a tandem repeat, otherwise you are not. Note that in this context n is the base differences between the ref and alt alleles	2012-04-06 16:07:46 -04:00
Mark DePristo	08fab49d30	Added function to get bases from the current base forward in the window in ReferenceContext	2012-04-06 16:07:46 -04:00
Ryan Poplin	c77104b815	Adding function call in HaplotypeCaller right before the VariantContext gets written out to disk which partitions all the reads by which allele gave the read the highest likelihood. This will allow variants to be annotated by the refactored VariantAnnotator. Uninformative reads are mapped to Allele.NO_CALL	2012-04-06 00:22:52 -04:00
Mauricio Carneiro	a19c27297f	continuing the BQSR triage... * fixed the loading of the new reduced size reports * reduced BQSR scala script memory to 2Gb * removed dcov parameter from BQSR scala script * fixed estimatedQReported calculation from -log10(pe) to -10log10(pe). updated md5's with the proper PHRED scaled EstimatedQReported	2012-04-05 14:34:15 -04:00
Eric Banks	3561056a9c	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-05 10:49:26 -04:00
Eric Banks	5c3ddec4c2	Large refactoring of the genotyping codebase. Deprecated several of the old classes that had the wrong allele ordering and made new better copies with the correct ordering; eventually we'll push the new ones into the place of the old ones but for now we'll give users a chance to update their code. Also, removed (or deprecated as needed) the genotype priors classes since we never use them and all they serve to do is make reading the code more complicated. I expect to finish this refactoring in GATK 1.7 (or 2.0?) so that should give Kristian ample time to update.	2012-04-05 10:49:08 -04:00
Mauricio Carneiro	7c3b3650bb	BQSR bug triage * fixed bug where some keys were using the same recal datum objects * fixed quantization qual calculations when combining multiple reports * fixed rounding error with empirical quality reported when combining reports * fixed combine routine in the gatk reports due to the primary keys being out of order * added auto-recalibration option to BQSR scala script * reduced the size of the recalibration report by ~15% * updated md5's	2012-04-05 09:32:18 -04:00
Eric Banks	2c956efa53	Minor fixups to GenotypeLikelihoods	2012-04-05 09:14:37 -04:00
Mauricio Carneiro	1e65474fec	Added utility to get the reference coordinate given the read coordinate	2012-04-05 09:04:20 -04:00
Guillermo del Angel	6913710e89	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-04 20:17:18 -04:00
Mark DePristo	76e4100d89	By default, IndelLengthHistogram won't collapse large events into the last bin, as it produces weird looking plots -- Updated integration tests as well	2012-04-04 18:48:03 -04:00
Guillermo del Angel	820216dc68	More pool caller cleanups: ove common duplicated code between Pool and Exact AF calculation models up to super-class to avoid duplication. TMP: Have pool genotypes include the GT field. Mostly because without genotypes we can't get the site-wide AF,AC annotations, but it's unwieldy because it makes the genotype columns very long, TBD final implementation	2012-04-04 16:23:10 -04:00
Ryan Poplin	bfad26353a	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-04 16:04:50 -04:00
Ryan Poplin	dda2173c66	Moved the Smith-Watermaning of haplotypes to earlier in the process so that alleles sent to genotyping would have the exact genomic sequence of the active region they represent. As a side effect cleaned up some edge case problems with variants, both real and false, which show up on the edges of active regions. Removed code that was replicated between the Haplotype class and ReadUtils. Finally figured out how to ensure that the indel calls coming out of the HC were left aligned.	2012-04-04 16:04:29 -04:00
Mark DePristo	fcdd65a0f4	Bugfix for IndelLengthHistogram -- Wasn't requiring the allele to actually be polymorphic in the samples, so it wasn't working correctly with the Sample strat.	2012-04-04 15:37:43 -04:00
Mark DePristo	1ccea866d8	VariantEval now includes -keepAC0 argument to include sites with alt alleles but AC 0 in analyses -- Updated EvalModules to work with new paramter -- adding test file for keepAC0 to public/testdata and integration tests	2012-04-04 15:37:12 -04:00
Eric Banks	9e32a975f8	Wow, symbolic alleles were all busted internally and this finally bubbled up after my previous commit. For some reason we were inconsistently forcing allele trimming/padding if one was present. Not anymore.	2012-04-04 13:47:59 -04:00
Eric Banks	337ff7887a	When constructing VariantContexts from symbolic alleles, check for the END tag in the INFO field; if present, set the stop position of the VC accordingly. Added integration test to ensure that this is working properly for use with -L intervals.	2012-04-04 10:57:05 -04:00
Guillermo del Angel	05d8400468	Fix up broken non-pool UG tests: GenotypeLikelihoods.calcNumLikelihoods now expects total # of alleles, not # of alt ones. Add doc to new function implementation. Add unit test for function. Add unit test for PoolGenotypeLikelihoods (not fully done yet)	2012-04-03 20:51:24 -04:00
Guillermo del Angel	5abb07da5d	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-03 17:00:45 -04:00
Christopher Hartl	a6837d31d4	Success! A fast and low-memory converter from VCF into a binary ped file. This is mostly so I don't have to listen to Pierre/Jason complain about how slow and inefficient plinkseq is at converting; or at transposting. This automatically writes to individual-major mode. It will eat up space on /tmp if you don't run with -Djava.io.tmpdir, so be careful if you use it.	2012-04-03 16:13:16 -04:00
Guillermo del Angel	63b1e737c6	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-04-03 15:43:50 -04:00
Guillermo del Angel	9e11b4f9a7	Major refactor/completion of new Pool Caller under UnifiedGenotyper framework. PoolAFCalculationModel implements new math to combine pools - correct, but still O(N^2) and not complete yet for multiallelics. Pool likelihoods are better encapsulated and kept in an internal hashmap from int[] -> double for space efficiency (likelihoods can be big for pool calls when in initial discovery mode with 4 alleles). Maybe need several iterations of optimization to make it runnable at large scale. Still need to correct function chooseMostLikelyAlternateAlleles before full runs can be produced.	2012-04-03 15:43:32 -04:00
Eric Banks	f9ce9962c4	Minor changes to verbose mode	2012-04-03 10:53:48 -04:00
Eric Banks	f6aa95685d	OutOfMemory exceptions are User Errors	2012-04-02 22:46:56 -04:00
Eric Banks	659b82e74d	Old -B syntax is long gone at this point. Safe to remove the warning.	2012-04-02 22:25:16 -04:00
Eric Banks	99d27ddcc4	Had some free time, so I unplugged extended events from the walkers. Now they exist only in LocusIteratorByState, but ReadProperties.generateExtendedEvents() always returns false so that block is never actually executed anymore. I don't want to touch LIBS because I think David is in there right now.	2012-04-02 14:27:36 -04:00
Mark DePristo	6b7a00061a	VariantsToTable now works with multiple input VCFs	2012-04-02 09:13:35 -04:00
Mark DePristo	fbbb8509ad	Final commits to VariantEval -- Molten now supports variableName and valueName so you don't have to use variable and value if you don't want to. -- Cleanup code, reorganize a bit more. -- Fix for broken integrationtests	2012-03-30 20:11:06 -04:00
Mark DePristo	4b45a2c99d	Final version of new VariantEval infrastructure. * WAY FASTER * -- 3x performance for multiple sample analysis with 1000 samples -- Analyzing 1MB of the ESP call set (3100 samples) takes 40 secs, compared to several minutes in the previous version -- According to JProfiler all of the runtime is now spent decoding genotypes, which will only get better when we move to BCF2 -- Remove the TableType system, as this was way too complex. No longer possible to embed what were effectively multiple tables in a single Evaluator. You now have to have 1 table per eval -- Replaced it with @Molten, which allows an evaluator to provide a single Map from variable -> value for analysis. IndelLengthHistogram is now a @Molten data type. GenotypeConcordance is also. -- No longer allow Evaluators to use private and protected variables at @DataPoints. You get an error if you do. -- Simplified entire IO system of VE. Refactored into VariantEvalReportWriter. -- Commented out GenotypePhasingEvaluator, as it uses the retired TableType -- Stratifications are all fully typed, so it's easy for GATKReports to format them. -- Removed old VE work around from GATKReportColumn -- General code cleanup throughout -- Updated integration tests	2012-03-30 15:31:56 -04:00
Mark DePristo	8c0718a7c9	Fixed missing import	2012-03-30 15:31:55 -04:00
Mark DePristo	097ed4ecc4	Memory usage optimizations and safety improvements to StratNode and StratificationManager -- Added memory and safety optimizations to StratNode and StratificationManager. Fresh, immutable Hashmaps are allocated for final data structures, so they exactly the correct size and cannot be changed by users. -- Added ability of a stratification to specify incompatible evaluation. The two strats using this are AC and Sample with VariantSummary, as this computes per-sample averages and so combining these results in an O(n^2) memory requirement. Added integration test to cover incompatible strats and evals	2012-03-30 15:31:55 -04:00
Mark DePristo	b335c22f6d	Fully refactored, mostly cleaned up version of VariantEval using StratificationManager	2012-03-30 15:31:55 -04:00
Mark DePristo	c8086a79e3	New StratificationManager based VariantEval passes unmodified integration tests -- Now needs cleanup and optimizations	2012-03-30 15:31:55 -04:00
Mark DePristo	d37f31e349	First version of VariantEval that runs (approximately correctly) with new StratificationManager	2012-03-30 15:31:54 -04:00
Mark DePristo	8971b54b21	Phase II of Stratification manager -- Renamed and reorganized infrastructure -- StratificationManager now a Map from List<Object> -> V. All key functions are implemented. Less commonly used TODO -- Ready for hookup to VE	2012-03-30 15:31:54 -04:00
Mark DePristo	9f1cd0ff66	Lots of new functionality for StratificationStates manager -- Really working according to unit tests -- A nCombination utils	2012-03-30 15:31:54 -04:00
Mark DePristo	a3d896d80e	Part I of creating a fast state space lookup for VE -- Created a unit tested tree mapping from a List<String> -> integer (StratificationStates). This class is the key infrastructure necessary to create a complete static mapping from all stratification combinations to an offset in a vector of EvalutionContexts for update in map. -- Minor code cleanup throughout VE (removing unused headers, for example)	2012-03-30 15:31:53 -04:00
Eric Banks	533c283783	Deprecating AlignmentContext.getExtendedEventPileup(). At this point the only walkers left with any relaiance on extended events are Guillermo's pooled code (he'll update soon) and the Pileup walker. David, I'll leave that last one for you (it should be easy). We can now officially rip the extended event code from the engine.	2012-03-30 10:37:14 -04:00
Eric Banks	6b49af253b	Removing dependence on extended events from the RealignerTargetCreator. Did some minor refactoring while I was in there.	2012-03-30 10:33:30 -04:00
Eric Banks	b467cd1dae	Removing dependence on extended events for the remaining Variant Annotator modules.	2012-03-30 09:05:26 -04:00
Eric Banks	b21889812d	Removing some more usages of extended events. Not done yet, but almost there.	2012-03-30 01:51:37 -04:00
Eric Banks	ad6ace2439	Resolving merge conflicts	2012-03-30 01:51:09 -04:00
Eric Banks	f4d4969f23	Don't ever return null for the list of GL models	2012-03-30 00:22:40 -04:00
Eric Banks	44ac49aa34	Removing dependencies in the annotations on extended events. Some refactoring involved in this.	2012-03-30 00:17:02 -04:00
Mauricio Carneiro	cbd21c6339	Nasty, nasty..... VariantEval is overly abusive of the GATKReport (lack of) spec. 1. It converts numeric values (longs, integers and doubles) to string before sending to the Report, then expects it to decipher that those were actually numbers. 2. Worse, the stratification modules somehow instead of sending the actual values to the report table, sends a string with the value "unknown" and then abuses the GATKReport spec to convert those "unknown" placeholder values with numbers. Then again, it expects the report to know those are numbers, not strings. Now that the GATKReport HAS specs, VariantEval needs to be overhauled to conform with that. In the meantime, I have added special ad-hoc treatment to these wrong contracts. It works, and the integration tests all passed without changing any MD5's, but right after Mark and Ryan commit their VariantEval refactors, I will step in to change the way it interacts with the GATKReport, so we can clean up the GATKReport. No wonder, the printing needed to be O(n^2).	2012-03-29 17:49:53 -04:00
Eric Banks	c2e27729c7	Renaming PileupElement.isBeforeDeletion() to PileupElement.isBeforeDeletedBase() so that it's more clear that it can still be true while inside a deletion. Added PileupElement.isBeforeDeletionStart() to cover the case that I want where we only trigger before the actual deletion event. Similarly for after a deletion. Updated counting code in ConsensusAlleleCounter accordingly.	2012-03-29 17:08:25 -04:00
Ryan Poplin	6da9571829	resolving merge conflicts.	2012-03-29 16:16:28 -04:00
Ryan Poplin	ca96544ed0	All the zero quality N bases in the solid reads are adding lots of extra paths in the assembly graph. We now require a minimum base quality for every base in the kmer before adding it to the graph. The large number of solid reads with unmapped mates was also triggering the active region traversal at every base. We now ignore that check for solid reads.	2012-03-29 16:14:29 -04:00
Eric Banks	e4469a83ee	First attempt at removing all traces of extended events from UG; integration tests are expected to fail.	2012-03-29 14:59:29 -04:00
Eric Banks	e61e162c81	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-03-29 12:33:13 -04:00
Mauricio Carneiro	cf364f26a0	Fixing alignment issue with the GATKReportColumn algorithm Numeric columns were being left-aligned when they should be right-aligned. Fixed it.	2012-03-29 12:28:49 -04:00
Mauricio Carneiro	f80bd4276a	fixed estimated Q reported calculation in the gatherer	2012-03-29 12:28:43 -04:00
Mauricio Carneiro	8a9fb514b6	simplifying GATKReportColumn constructor logic	2012-03-29 12:28:37 -04:00
Eric Banks	e861106398	Accidentally erased important line	2012-03-29 11:08:54 -04:00
Eric Banks	e4a225ed09	Move the code to subset a Variant Context to fewer alleles (including restructuring the PLs appropriately) into VariantContextUtils where it can be used generally.	2012-03-29 11:07:37 -04:00
Guillermo del Angel	c9c3f6b0fc	Minor UG Engine refactoring/cleanup: instead of passing in the # of samples separately from sample set, pass in ploidy instead and compute # of chromosomes internally - will help later on with code clarity	2012-03-29 11:05:42 -04:00
Ryan Poplin	9684a2efb0	HaplotypeCaller: Variants found on the same haplotype are now written out with phased genotypes. There are serious eval issues with MNPs so disabling them for now.	2012-03-29 09:41:29 -04:00
Guillermo del Angel	250adca350	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-03-28 21:01:49 -04:00
Guillermo del Angel	e0ab4e4b30	Refactoring so that ConsensusAlleleCounter can use regular pileups and can operate correctly. This involved adding utility functions to ReadBackedPileup to count # of insertions/deletions right after current position. Added unit test for IndelGenotypeLikelihoods, esp. ConsensusAlleleCounter logic	2012-03-28 21:01:31 -04:00
Mauricio Carneiro	8f0e9d74ce	GATKReportTable output refactor writing out a GATKReportTable was O(n^2)!!!!! New implementation is O(n). What a difference, when N = 2^16...	2012-03-28 17:19:12 -04:00
Guillermo del Angel	62ee31afba	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-03-28 16:00:38 -04:00
Guillermo del Angel	1eee9d512d	Make computeConsensusAlleles protected inside IndelGenotypeLikelihoodsCalculationModel so we can use it in unit tests, b) make ConsensusAlleleCounter work if no extended event pileup is present (necessary for ext. event removal)	2012-03-28 15:41:39 -04:00
Mauricio Carneiro	bb36cd4adf	Quick fixes to BQSRGatherer and GATKReportTable * when gathering, be aware that some keys will be missing from some tables. * when a gatktable has no elements, it should still output the header so we know it had no records	2012-03-28 09:07:54 -04:00
Roger Zurawicki	63cf7ec7ec	Added more primitives to GATK Report Column Type - The Integer column type now accepts byte and shorts - Updated Unit Tests and added a new testParse() test Signed-off-by: Mauricio Carneiro <carneiro@broadinstitute.org>	2012-03-28 09:07:54 -04:00
Guillermo del Angel	08f7d47d7c	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-03-28 07:42:09 -04:00
Mark DePristo	12aa72f200	Merged bug fix from Stable into Unstable	2012-03-27 22:43:00 -04:00
Mark DePristo	979a84a252	Bugfix for thread unsafe PL cache -- See https://getsatisfaction.com/gsa/topics/unifiedgenotyper_error_indel?utm_content=topic_link&utm_medium=email&utm_source=new_topic -- Solution is to use a fixed cache that's never updated on the fly. My changes limit us to having no more than 500 alleles at a site, which I hope is ok but easy enough to up to a ridiculously large number.	2012-03-27 22:42:30 -04:00
Guillermo del Angel	8f34412fb8	First Pool Caller exact model: silly straightforward math implementation of biallelic pool caller exact likelihood model, no attempt and any smartness or optimization, no support yet for generalized multiallelic form, just hooking up for testing	2012-03-27 20:59:44 -04:00
Guillermo del Angel	ed322bd73f	Fix again merge issues	2012-03-27 15:03:13 -04:00
Guillermo del Angel	b4a7c0d98d	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-03-27 15:01:03 -04:00
Guillermo del Angel	343a061b1c	Fix merge issues when incorporating new AF calculations changes	2012-03-27 15:00:44 -04:00
Mauricio Carneiro	1b75663178	BQSR Gatherer implementation and integration tests * restructured the hash tables into one class (RecalibrationReport) that has all the functionality for the different tables and key managers * optmized empirical qual calculation when merging recalibration reports * centralized the quality score quantization functionalities * unified the creating/loading of all the key manager/hash table structures. * added unit tests for the gatherer (disabled because gatk report needs to be sorted for automated testing) * added integration tests for BQSR and on-the-fly recalibration	2012-03-27 13:50:22 -05:00
Ryan Poplin	5dbd3625cd	Initial algorithm for choosing best alternate haplotypes to genotype based on the likelihoods from all samples instead of choosing for each sample independently. Simple tradeoff of penalty for increasing model complexity and likelihood of the data.	2012-03-27 13:38:52 -04:00
Eric Banks	c112e0824a	I was adding verbose output to the Pileup output for a one-off and decided that I might as well commit it as an option. Updated deprecated calls while I was in there.	2012-03-27 11:09:03 -05:00
Mark DePristo	a638996fe2	Cleanup of VariantEval, diatribe about performance problems with StateKey -- Minor refactoring of state key iteration in VEW.map to make the dependencies more clear -- Long discussion about the performance problems with StateKey, and how to fix it, which I have run out of time to address before ESP meeting.	2012-03-27 11:56:24 -04:00
Mark DePristo	679bb03014	Simple utility function for converting an Iterable<T> to Collection<T>	2012-03-27 11:54:58 -04:00
Mark DePristo	1f5f737c8b	Optimizing the GATKReportTable.write -- Better iteration, caching of strings, better printf calls, to improve the writing performance of GATKReportTables	2012-03-27 11:54:35 -04:00
Mark DePristo	913c8b231f	Fix ErrorRatePerCycle to overload equals and hashcode -- Fixes failing integration tests	2012-03-27 10:35:32 -04:00
Eric Banks	c07a577ba3	Significant restructuring of the Exact model, as discussed within the dev group last week. There is no more marginalizing over alternate alleles, and we now keep track of the MLE and MAP. Important notes: 1) integration tests change because the previous marginalization wasn't done correctly (as pointed out by Guillermo) and our confidences were too high for many multi-allelic sites; 2) there is a major TO-DO item that needs to be discussed within the dev group (so they should expect a follow up email); 3) this code is still in flux as I am awaiting feedback from Ryan now on its performance with the Haplotype Caller (the good news, Ryan, is that we recover that site that we were losing previously).	2012-03-27 00:27:44 -05:00
Mark DePristo	34ea443cdb	Better algorithm for choosing which indel alleles are present in samples -- The previous approach (requiring > 5 copies among all reads) is breaking down in many samples (>1000) just from sequencing errors. -- This breakdown is producing spurious clustered indels (lots of these!) around real common indels -- The new approach requires >X% of reads in a sample to carry an indel of any type (no allele matching) to be including in the counting towards 5. This actually makes sense in that if you have enough data we expect most reads to have the indel, but the allele might be wrong because of alignment, etc. If you have very few reads, then the threshold is crossed with any indel containing read, and it's counted. -- As far as I can tell this is the right thing to do in general. We'll make another call set in ESP and see how it works at scale. -- Added integration tests to ensure that the system is behaving as I expect on the site I developed the code on from ESP	2012-03-26 16:28:49 -04:00
Mark DePristo	11b6fd990a	GATKReportColumn optimizations -- Was TreeMap even though the sorting wasn't used. Replaced with LinkedHashMap.	2012-03-26 16:28:49 -04:00
Mark DePristo	6be5e82860	VariantEval scalability optimizations -- StateKey no longer extends TreeMap. It's now a final immutable data structure that caches it's toString and hashcode values. TODO optimizations to entirely remove the TreeMap and just store the HashMap for performance and use the tree for the sorted tostring function. -- NewEvaluationContext has a method makeStateKey() that contains all of the functionality that once was spread around VEUtils -- AnalysisModuleScanner uses an annotationCache to speed up the reflections getAnnotations() call when invoked over and over on the same objects. Still expensive to convert each field to a string for the cache, but the only way around that is a complete refactoring of the toTransversalDone of VE -- VariantEvaluator base class has a cached getSimpleName() function -- VEUtils: general cleanup due to refactoring of StateKey -- VEWalker: much better iteration of map data structures. If you need access to iterate over all key/value pairs use the Map.Entry construct with entrySet. This is far better than iterating over the keys and calling get() on each key.	2012-03-26 16:28:48 -04:00
Guillermo del Angel	1c424c0daf	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-03-26 15:15:50 -04:00
Ryan Poplin	019145175b	Major optimizations to graph construction through better use of built in graph.containsVertex and vertex.equals methods. Minor optimizations to MathUtils.approximateLog10SumLog10 method	2012-03-26 11:32:44 -04:00
Ryan Poplin	1fa66f76c9	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-03-25 23:04:47 -04:00
Guillermo del Angel	ce617b2dfc	Bug fix to previous UnifiedGenotyperEngine refactoring, removed debug code	2012-03-25 10:20:21 -04:00
Guillermo del Angel	db54c2625f	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-03-25 09:53:35 -04:00
Guillermo del Angel	deb4586559	Next intermediate commit for new pool caller structure: a) Bug fixes in pool GL computation. Now, correct GL's are returned per each pool to the UG engine. Work still needs to be done in redoing interface with exact model. b) Added unit tests for new MathUtils dot product and logDotProduct functions. c) Refactorings of UnifiedGentotyperEngine since N (size of prior/posterior arrays) is no longer necessarily nSamples+1 but, in general, nSamplesPerPool*nPools+1	2012-03-24 21:49:43 -04:00
Mark DePristo	b063bcd38d	Removing update0 support in VariantEval -- Now the only use for update0, calculating the number of processed loci, is centrally tracked in the walker itself not the evaluations. -- This allows us to avoid calling update0 are every genomic base in 100ks of evaluates when there are a lot of stratifications. -- No need to modify the integration tests, this optimization doesn't change the result of the calculation	2012-03-23 21:02:21 -04:00
Mauricio Carneiro	0509d316d9	More information in the recalibration report * added empirical quality counts to allow quantization during on-the-fly recalibration to any level * added number of observations and errors to all tables to enable plotting of all covariates	2012-03-23 16:15:19 -04:00
Mauricio Carneiro	9f74969e3a	BQSR with GATKReport implementation * restructured BQSR to report recalibrated tables. * implemented empirical quality calculation to the BQSR stage (instead of on-the-fly recalibration) * linked quality score quantization to the BQSR stage, outputting a quantization histogram * included the arguments used in BQSR to the GATK Report * included all three tables (RG, QUAL and COVARIATES) to the GATK Report with empirical qualities On-the-fly recalibration with GATK Report * loads all tables from the GATKReport using existing infrastructure (with minor updates) * implemented initialiazation of the covariates using BQSR's argument list * reduced memory usage significantly by loading only the empirical quality and estimated quality reported for each bit set key * applied quality quantization to the base recalibration * excluded low quality bases from on-the-fly recalibration for mismatches, insertions or deletions	2012-03-23 15:42:32 -04:00
Mauricio Carneiro	f421062b55	Updated read group covariate to use sample.lane instead of the id Added Unit test.	2012-03-23 15:24:07 -04:00
Mauricio Carneiro	539da9e3e1	Fixing GATKReport exception handling when loading a report * allowing tables with no description to go through * GATKReportTable should be more lenient with the format requirements (added to-dos for roger)	2012-03-23 15:23:13 -04:00
Eric Banks	2511839068	Merged bug fix from Stable into Unstable	2012-03-23 13:51:33 -04:00
Eric Banks	d3f2bc4361	Pre-allocate 10 alt alleles worth of PLs in the cache for efficiency. This effectively means that we never need to re-allocate the cache in the future because we can't ever really handle that many alt alleles.	2012-03-23 13:51:00 -04:00
Mark DePristo	e4ec90cfce	Merged bug fix from Stable into Unstable	2012-03-23 11:27:34 -04:00
Mark DePristo	ff26f2bf68	HierarchicalMicroScheduler no longer attempts to wrap exceptions -- This behavior, which isn't obviously valuable at all, continued to grab and rethrow exceptions in the HMS that, if run without NT, would show up as more meaningful errors. Now HMS simply checks whether the throwable it received on error was a RuntimeException. If so, it is stored and rethrow without wrapping later. If it isn't, only in this case is the exception wrapped in a ReviewedStingException. -- Added a QC walker ErrorThrowingWalker that will throw a UserException, ReviewedStingException, and NullPointerException from map as specified on the command line -- Added IT that ensures that all three types are thrown properly (i.e., you catch a NullPointerException when you ask for one to be thrown) with and without threading enabled. -- I believe this will finally put to rest all of these annoying HMS captures.	2012-03-23 11:27:21 -04:00
Ryan Poplin	9d22471b79	Merged bug fix from Stable into Unstable	2012-03-23 10:48:34 -04:00
Ryan Poplin	ab288354e9	Better error message for malformed input recal file.	2012-03-23 10:47:01 -04:00
Mark DePristo	fee8d86f63	VariantEval optimization -- Use a LinkedHashMap not a TreeMap so iteration is faster. -- Note that with a lot of stratifications the update0 is taking up a lot of time. For example, with 822 samples and functional class and sample on there are 100K contexts and 30% of the runtime is just in the update0 call	2012-03-22 22:13:24 -04:00
Mark DePristo	6df96644d9	Unified, standard IndelSummary metrics for VariantEval -- Now you always get SNP and indel metrics with VariantEval! -- Includes Number of SNPs, Number of singleton SNPs, Number of Indels, Number of singleton Indels, Percent of indel sites that are multi-allelic, SNP to indel ratio, Singleton SNP to indel ratio, Indel novelty rate, 1 to 2 bp indel ratio, 1 to 3 bp indel ratio, 2 to 3 bp indel ratio, 1 and 2 to 3 bp indel ratio, Frameshift percent, Insertion to deletion ratio, Insertion to deletion ratio for 1 bp events, Number of indels in protein-coding regions labeled as frameshift, Number of indels in protein-coding regions not labeled as frameshift, Het to hom ratio for SNPs, Het to hom ratio for indels, a Histogram of indel lengths, Number of large (>10 bp) deletions, Number of large (>10 bp) insertions, Ratio of large (>10 bp) insertions to deletions -- Updated VE integration tests as appropriate	2012-03-22 21:24:37 -04:00
Mark DePristo	bcf80cc7b3	Cleanup in VariantEval. Example of molten VariantEval output -- Moved a variety of useful formatting routines for ratios, percentages, etc, into VariantEvalator.java so everyone can share. Code updated to use these routines where appropriate -- Added variantWasSingleton() to VariantEvaluator, which can be used to determine if a site, even after subsetting to specific samples, was a singleton in the original full VCF -- TableType, which used to be an interface, is now an abstract class, allowing us to implement some generally functionality and avoid duplication. -- This included creating a getRowName() function that used to be hardcoded as "row" but how can be overridden. -- #### This allows us implement molten tables, which are vastly easier to use than multi-row data sets. See IndelHistogram class (in later commit) for example of molten VE output	2012-03-22 21:24:37 -04:00
Mark DePristo	9ddd5aec93	More eval modules being removed from VariantEval -- IndelStatistics is superceded by IndelStatistics	2012-03-22 21:24:36 -04:00
Mark DePristo	bd5b6d1aba	Remove no longer in use Eval modules from VariantEval -- No more IndelLengthHistogram (superceded by IndelSummary in subsequent commit) -- No more SamplePreviousGenotypes or PhaseStats -- No more MultiallelicAFs	2012-03-22 21:24:36 -04:00
Menachem Fromer	7faa9938b1	Merge branch 'master' of ssh://copper.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-03-22 17:43:44 -04:00
Menachem Fromer	b9b9219ac7	Added respectPhaseInInput flag to RBP and integration tests	2012-03-22 17:40:21 -04:00
Guillermo del Angel	f198cec5e2	Temp commit: new structure for pool caller, now all work is in the same framework as in UG. There's a new genotype calculation model, PoolGenotypeCalculationModel, that does all the work and plugs into UnifiedGenotyperEngine. A new AF module for pools is upcoming. Old pool caller will be removed once all work is migrated	2012-03-22 15:46:39 -04:00
Menachem Fromer	1dfaacfeb5	Check for consistency of the BAM and VCF sample names, with a command line disable to throw if you know what you are doing	2012-03-22 12:40:15 -04:00
Guillermo del Angel	b02ef95bcf	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-03-22 12:14:12 -04:00
Guillermo del Angel	92676c63ca	Make constructor of IndelGenotypeLikelihoodsCalculationModel public so it can be used in unit tests	2012-03-22 12:13:59 -04:00
Guillermo del Angel	58965d6a6e	Merged bug fix from Stable into Unstable	2012-03-22 11:04:11 -04:00
Guillermo del Angel	b8cd959461	Potential corner condition bug fix: protect against null pointer exceptions when computing consensus indel bases when UG is discovering alt alleles. If an alt allele has non-standard bases, skip allele gracefully instead of adding null object into list	2012-03-22 10:06:22 -04:00
Ryan Poplin	a29fc6311a	New debug option to output the assembly graph in dot format. Merge nodes in assembly graph when possible.	2012-03-21 15:48:55 -04:00
Eric Banks	8c09ff9459	Merged bug fix from Stable into Unstable	2012-03-21 12:44:43 -04:00
Eric Banks	58245bfa2f	Bug fix: check to see whether there's a BasePileup before asking for one.	2012-03-21 12:44:09 -04:00
Eric Banks	07c3bd32b3	Bug fix: merge NO_VARIATION records with those of another type. The sad part is that this WAS covered by integration tests but someone updated the MD5s without actually paying attention...	2012-03-21 12:42:13 -04:00
Eric Banks	dcf2fa361d	Minor cleanup	2012-03-21 12:14:31 -04:00
Eric Banks	ab1c48745b	Need to catch RuntimeExceptions coming out of Picard too so that they show up as UserErrors (some BAM errors are thrown as REs).	2012-03-21 12:13:52 -04:00
Ryan Poplin	9e10779fa7	Caching log calculations cut the non-Map runtime of HaplotypeCaller in half. Moved the qual log cache used in HC and PairHMM into a common place and added unit tests.	2012-03-21 08:45:42 -04:00
Mauricio Carneiro	0e93cf5297	Taking care of bad cigars in the GATK * fixed BadCigarFilter to filter out reads starting/ending in deletion and that have adjacent I/D events. * added Unit tests for BadCigarFilter * updated all exceptions in LocusIteratorByState to tell the user that he can instead run with -rf BadCigar * added the BadCigar filter to ReduceReads and RealignTargetCreator (if your walker blows up with these malformed reads, you may want to add it too)	2012-03-20 14:32:57 -04:00
Eric Banks	5e79046c98	Minor change but I realized from Mark's commit that the code I stole it from was flawed	2012-03-20 08:55:56 -04:00
Eric Banks	ade1971581	Since we allow any generic header types, there's no longer any reason to check for supported types	2012-03-20 00:12:17 -04:00
Eric Banks	2324c5a74f	Simplified the interface for simple VCF header lines by making the VCFSimpleHeaderLine not abstract anymore - now any arbitrary header line with an ID (e.g. the contig and ALT lines) can be part of this class without having to define new classes. Also, renamed the 'named' header line to 'id' since that's more accurate.	2012-03-19 21:29:24 -04:00
Roger Zurawicki	7afb333811	GATK Report code cleanup - Updated the documentation on the code - Made the table.write() method private and updated necessary files. - Added a constructor to GATKReport that takes GATKReportTables - Optimized my code Signed-off-by: Mauricio Carneiro <carneiro@broadinstitute.org>	2012-03-19 11:53:57 -04:00
Mauricio Carneiro	0d4ea30d6d	Updating the BQSR Gatherer to the new file format This is important for quick turnaround in the analysis cycle of the new covariates. Also added a dummy unit test that doesn't really test anything (disabled), but helps in debugging.	2012-03-19 09:02:27 -04:00
Eric Banks	9223e451a3	Merged bug fix from Stable into Unstable	2012-03-18 00:54:19 -04:00
Eric Banks	344a938a70	When checking to make sure that we have cached enough data in the PL array, use the converted index value since that's what will be used as an index into the array.	2012-03-18 00:36:30 -04:00
Eric Banks	be9e48ba29	Merged bug fix from Stable into Unstable	2012-03-16 14:33:53 -04:00
Mauricio Carneiro	ec4a870a0f	Added @PG tag to ReduceReads Pulled out the functionality from Indel Realigner and Table Recalibrator into Utils.setupWriter to make everyone else's life's easier if they want to include the PG tag in their walkers.	2012-03-16 14:09:07 -04:00
Mauricio Carneiro	3bfca0ccfd	BitSet implementation of the on-the-fly recalibration using the CSV format file. Infrastructure: * Added static interface to all different clipping algorithms of low quality tail clipping * Added reverse direction pileup element event lookup (indels) to the PileupElement and LocusIteratorByState * Complete refactor of the KeyManager. Much cleaner implementation that handles keys with no optional covariates (necessary for on-the-fly recalibration) * EventType is now an independent enum with added capabilities. All functionality is now centralized. BQSR and RecalibrateBases: * On-the-fly recalibration is now generic and uses the same bit set structure as BQSR for a reduced memory footprint * Refactored the object creation to take advantage of the compact key structure * Replaced nested hash maps with single hash maps indexed by bitsets * Eliminated low quality tails from the context covariate (using ReadClipper's write N's algorithm). * Excluded contexts with N's from the output file. * Fixed cycle covariate for discrete platforms (need to check flow cycle platforms now!) * Redfined error for indels to look at the previous base in negative strand reads (using new PE functionality) * Added the covariate ID (for optional covariates) to the output for disambiguation purposes * Refactored CovariateKeySet -- eventType functionality is now handled by the EventType enum. * Reduced memory usage of the BQSR script to 4 Tests: * Refactored BQSRKeyManagerUnitTest to handle the new implementation of the key manager * Added tests for keys without optional covariates * Added tests for on-the-fly recalibration (but more tests are necessary)	2012-03-16 13:02:15 -04:00
Mauricio Carneiro	ca11ab39e7	BitSets keys to lower BQSR's memory footprint Infrastructure: * Generic BitSet implementation with any precision (up to long) * Two's complement implementation of the bit set handles negative numbers (cycle covariate) * Memoized implementation of the BitSet utils for better performance. * All exponents are now calculated with bit shifts, fixing numerical precision issues with the double Math.pow. * Replace log/sqrt with bitwise logic to get rid of numerical issues BQSR: * All covariates output BitSets and have the functionality to decode them back into Object values. * Covariates are responsible for determining the size of the key they will use (number of bits). * Generalized KeyManager implementation combines any arbitrary number of covariates into one bitset key with event type * No more NestedHashMaps. Single key system now fits in one hash to reduce hash table objects overhead Tests: * Unit tests added to every method of BitSetUtils * Unit tests added to the generalized key system infrastructure of BQSRv2 (KeyManager) * Unit tests added to the cycle and context covariates (will add unit tests to all covariates)	2012-03-16 13:01:48 -04:00
Eric Banks	dce6b91f7d	Add a conversion from the deprecated PL ordering to the new one. We need this for the DiploidSNPGenotypeLikelihoods which still use the old ordering. My intention is for this to be a temporary patch, but changing the ordering in DiploidSNPGenotypeLikelihoods is not appriopriate for committing to stable as it will break all of the external tools (e.g. MuTec) that are built on top of the class. We will have to talk to e.g. Kristian to see how disruptive this will be. Added unit tests to the GL conversions and indexing.	2012-03-16 11:14:37 -04:00
Eric Banks	41068b6985	The commit constitutes a major refactoring of the UG as far as the genotype likelihoods are concerned. I hate to do this in stable, but the VCFs currently being produced by the UG are totally busted. I am trying to make just the necessary changes in stable, doing everything else in unstable later. Now all GL calculations are unified into the GenotypeLikelihoods class - please try and use this functionality from now on instead of duplicating the code.	2012-03-15 16:08:58 -04:00
Ryan Poplin	0c6b34e9df	Fixing a bug identified by the ActivityProfile unit tests	2012-03-15 14:24:30 -04:00
Ryan Poplin	252b830aa8	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-03-15 11:56:04 -04:00
Ryan Poplin	1429ddcf55	Adding contracts and unit tests for HaplotypeCaller LikelihoodCalculationEngine	2012-03-14 21:25:43 -04:00
Mark DePristo	7c5cdb51c2	UnitTests for ActivityProfile and minor ART cleanup -- TODO for ryan -- there are bugs in ActivityProfile code that I cannot fix right now :-( -- UnitTesting framework for ActivityProfile -- needs to be expanded -- Minor helper functions for ActiveRegion to help with unit tests	2012-03-14 17:26:37 -04:00
Mark DePristo	e440c9be98	Clean up logic for adding reads to ART cache -- No longer has duplicate code	2012-03-14 17:26:37 -04:00
Mark DePristo	5bcb5c7433	Preliminary refactoring of ART -- Refactored ART into clearer, simpler procedures. Attempted to merge shared code into utility classes. -- Added some docs -- Created a new, testable ActivityProfile that represents as a class the probability of a base being active or inactive -- Separated band-pass filtering from creation of active regions. Now you can band pass filter a profile to make another profile, and then that is explicitly converted to active regions -- Misc. utility functions in ActiveRegionWalker such as hasPresetActiveRegions() -- Many TODOs in ActivityProfile.	2012-03-14 17:26:37 -04:00
Ryan Poplin	1da8928407	HC GenotypingEngine marginalizes over haplotypes when outputing events that were found on a subset of the called haplotypes.	2012-03-14 15:22:21 -04:00
Guillermo del Angel	eca055ccad	Add option in ValidationAmplicons to only output SNPs and INDELs, ignoring complex variants (or SVs, etc.)	2012-03-14 14:26:40 -04:00
Eric Banks	f7c2c818fe	Exact model memory optimization: instead of having a later matrix column pull in data from earlier ones (requiring us to keep them around until all dependencies are hit), the earlier columns push data into their dependents immediately and then are removed. This does trade off speed a little bit (because we need to call approximateLog10Sum each time we add to a dependent instead of once in an array at the end). Note that this commit would normally not get pushed into stable, but I'm about to make a very disruptive push into stable that would make merging this from unstable a nightmare.	2012-03-14 14:02:36 -04:00
Mark DePristo	6a40ca6bec	Merged bug fix from Stable into Unstable	2012-03-14 12:19:33 -04:00
Mark DePristo	bb2c10b785	Capture the class of the exception in GATKRunReport -- As suggested by David.	2012-03-14 12:16:22 -04:00
Ryan Poplin	78a4e7e45e	Major restructuring of HaplotypeCaller's LikelihoodCalculationEngine and GenotypingEngine. We no longer create an ugly event dictionary and genotype events found on haplotypes independently by finding the haplotype with the max likelihood. Lots of code has been rewritten to be much cleaner.	2012-03-14 12:05:05 -04:00
Eric Banks	77243d0df1	Splitting up the MultiallelicSummary module into the standard part for use by all and the dev piece used just by me	2012-03-13 16:31:51 -04:00
Eric Banks	568a1362f5	Splitting up the MultiallelicSummary module into the standard part for use by all and the dev piece used just by me	2012-03-13 16:19:15 -04:00
Eric Banks	5d7c761784	Merged bug fix from Stable into Unstable	2012-03-13 11:01:03 -04:00
Eric Banks	5200f7f919	When creating a synthetic VC based on the passed in alleles, set the reference base for indel.	2012-03-13 10:59:58 -04:00
Eric Banks	1675bd4dd7	When creating a synthetic VC based on the passed in alleles, set the length correctly.	2012-03-13 10:55:52 -04:00
Roger Zurawicki	7887a06703	GATKReport v1.0 GATKReport format changes: - All non-data header lines are preceeded with a single pound ( #:) - Every report now has a report header containing the version number and number of tables - Every table has two lines of table header: The first explains the size of the table and the data types of each column, the second contains the table name and description. - This new format will allow reports in the future to be gatherable. - Changed the header format to include an end-of-line string ":;" Added features: - Simplified GATK Reports: The constructor for a simplified GATK Report. Simplified GATK report are designed for reports that do not need the advanced functionality of a full GATK Report. A simple GATK Report consists of: - A single table - No primary key ( it is hidden ) Optional: - Only untyped columns. As long as the data is an Object, it will be accepted. - Default column values being empty strings. Limitations: - A simple GATK report cannot contain multiple tables. - It cannot contain typed columns, which prevents arithmetic gathering. - Added a constructor to generate simplified GATK reports. - Added a method to easily add data to simple GATK reports. - Upgraded the input parser take advantage of the new file format (v1). - Added the GATKReportGatherer, more usability cmoing in next versionof GATK Report. Curently, it can only add rows from one table to another. Added private methods in GATKReport to combine Tables and Reports, It is very conservative and will only gather if the table columns, as well as everything else matches. At the column level, it uses the (redundant) row ids to add new rows. It will throw an exception if it is overwriting data. - Made some GATKReport methods public, and added more setters and getters. - Added method that compares formats of two GATKReports, and added an equals method to verify all data inside. - The gsalib for R now supports reading GATKReport v1 files in addition to legacy formats (v0.) - Added a GATKReportDataType enum to give column a certain data type. This must be specified when making a gatherable report. This enum contains several methods including a reverse lookup map. - Added a data type field in GATKColumn, when a type is not specified, the unknown type is used. Unknown types should not be gathered. Test changes: - Updated Unit Tests for GATK Report v1. Added a test for the gatherer. Left one test disabled while we transition from v0 to v1. - Updated the MD5 hashes in integration tests throughout the GATK. Other changes: - Added the gatherer functions to CoverageByRG - Also added the scatterCount parameter in the Interval Coverage script - Dropped support for reading in legacy GATKReport formats ( v0.) - Updated VariantEvalWalker to work with GATK Report v1, added a format String to all applicable DataPoints. - Rewrote the read file method for GATK report files. - Optimized the equals methods within GATKReport. The protected functions should only be called by the GATKReport methods. Signed-off-by: Mauricio Carneiro <carneiro@broadinstitute.org>	2012-03-12 23:09:19 -04:00
Eric Banks	10995d349e	Fix old error message	2012-03-12 22:56:08 -04:00
Eric Banks	2314787767	Generalizing to avoid JDK 1.7 incompatibilities	2012-03-12 22:50:59 -04:00
Ryan Poplin	03223029e3	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-03-12 09:42:37 -04:00
Eric Banks	b4749757f8	Fixes for SLOD: 1) didn't work properly for multi-allelics (randomly chose an allele, possibly one that wasn't genotyped in the full context); 2) in cases when there were more alt alleles than the max allowed and the user is calculating SB, we would recompute the best alt alleles(s); 3) for some reason, we were recomputing the LOD for the full context when we'd already done that. Given that this passes integration tests on my end, this should be the last commit before the release.	2012-03-12 01:07:07 -04:00
Ryan Poplin	2836c161ee	Moving trimToVariableRegion out of reduced reads and into a public static ReadClipper function. HaplotypeCaller clips reads to the active region boundries before passing to the HMM. The philosophy of the HC is moving towards genotyping the entire haplotype sequence contained within the active region as a single allele.	2012-03-11 14:45:59 -04:00
Mark DePristo	1ee46e5c06	Collect only the bare essentials in the GATKRunReport Now looks like: <GATK-run-report> <id>D7D31ULwTSxlAwnEOSmW6Z4PawXwMxEz</id> <start-time>2012/03/10 20.21.19</start-time> <end-time>2012/03/10 20.21.19</end-time> <run-time>0</run-time> <walker-name>CountReads</walker-name> <svn-version>1.4-483-g63ecdb2</svn-version> <total-memory>85000192</total-memory> <max-memory>129957888</max-memory> <user-name>depristo</user-name> <host-name>10.0.1.10</host-name> <java>Apple Inc.-1.6.0_26</java> <machine>Mac OS X-x86_64</machine> <iterations>105</iterations> </GATK-run-report> No longer capturing command line or directory information, to minimize people's concerns with phone home and privacy	2012-03-10 20:27:14 -05:00
Mark DePristo	3ba2e5667c	CalibrateGenotypesLikelihoods include pOfDGivenD now	2012-03-09 16:00:07 -05:00
David Roazen	91d10431d3	BAMScheduler: detect contigs from the interval list that are not in the merged BAM header's sequence dictionary This is a quick-and-dirty patch for the null pointer error Mauricio reported earlier. Later on we might want to address in a more general way the fact that we validate user intervals against the reference but not against the merged BAM header produced by the engine at runtime.	2012-03-09 15:20:16 -05:00
David Roazen	bc65f6326f	Detect incomplete reads from BAM schedule file in BAMSchedule before they become buffer underflows This fix is similar, but distinct from the earlier fix to GATKBAMIndex. If we fail to read in a complete 3-integer bin header from the BAM schedule file that the engine has written, throw a ReviewedStingException (since this is our problem, not the user's) rather than allowing a cryptic buffer underflow error to occur. Note that this change does not fix the underlying problem in the engine, if there is one (there may be an as-yet-undetected bug in the code that writes the bam schedule). It will just make it easier for us to identify what's going wrong in the future.	2012-03-09 12:33:48 -05:00
David Roazen	32dee7ed9b	Avoid buffer underflow in GATKBAMIndex by detecting premature EOF in BAM indices GATKBAMIndex would allow an extremely confusing BufferUnderflowException to be thrown when a BAM index file was truncated or corrupt. Now, a UserException is thrown in this situation instructing the user to re-index the BAM. Added a unit test for this case as well.	2012-03-08 15:30:44 -05:00
Guillermo del Angel	c04853eae6	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-03-08 12:30:04 -05:00
Guillermo del Angel	858acf8616	Hidden mode in ValidationAmplicons to support ILMN output format (same as Sequenom, with just shuffled columns)	2012-03-08 12:29:44 -05:00
Andrey Sivachenko	56f074b520	docs updated	2012-03-07 18:47:15 -05:00
Andrey Sivachenko	117ea605ac	Merge branch 'master' of ssh://gsa1/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-03-07 18:35:07 -05:00
Andrey Sivachenko	497a1b059e	transition to JEXL completed, old parameters setting individual cutoffs now deprecated	2012-03-07 18:34:11 -05:00
Andrey Sivachenko	fbd2f04a04	JEXL support added; intermediate commit, not yet functional	2012-03-07 17:29:42 -05:00
Mark DePristo	0376d73ece	Improved, public version of ErrorRateByCycle -- A cleaner table output (molten). For those interested in seeing how this can be done with GATKReports look here for a nice clean example -- Integration tests -- Minor improvements to GATKReportTable with methods to getPrimaryKeys	2012-03-07 13:10:08 -05:00
Christopher Hartl	a6a8fc0521	Merge branch 'master' of ssh://ni.broadinstitute.org/humgen/gsa-scr1/chartl/dev/unstable	2012-03-07 10:05:43 -05:00
Mark DePristo	569be953b9	Bugfix for VariantEval -- We weren't properly handling the case where a site had both a SNP and indel in both eval and comp. These would naturally pair off as SNP x SNP and INDEL x INDEL in eval, but we'd still invoke update2 with (null, SNP) and (null, INDEL) resulting most conspicously as incorrect false negatives in the validation report. -- Updating misc. integrationtests, as the counting of comps (in particular for dbSNP) was inflated because of this effect.	2012-03-06 16:56:59 -05:00
Christopher Hartl	67def6acc8	Merge branch 'master' of ssh://ni.broadinstitute.org/humgen/gsa-scr1/chartl/dev/unstable	2012-03-06 14:23:14 -05:00
Christopher Hartl	20c1fbaf0f	Fixing a merge (turning off downsampling on DoC)	2012-03-06 14:22:45 -05:00
David Roazen	0702ee1587	Public-key authorization scheme to restrict use of NO_ET -Running the GATK with the -et NO_ET or -et STDOUT options now requires a key issued by us. Our reasons for doing this, and the procedure for our users to request keys, are documented here: http://www.broadinstitute.org/gsa/wiki/index.php/Phone_home -A GATK user key is an email address plus a cryptographic signature signed using our private key, all wrapped in a GZIP container. User keys are validated using the public key we now distribute with the GATK. Our private key is kept in a secure location. -Keys are cryptographically secure in that valid keys definitely came from us and keys cannot be fabricated, however keys are not "copy-protected" in any way. -Includes private, standalone utilities to create a new GATK user key (GenerateGATKUserKey) and to create a new master public/private key pair (GenerateKeyPair). Usage of these tools will be documented on the internal wiki shortly. -Comprehensive unit/integration tests, including tests to ensure the continued integrity of the GATK master public/private key pair. -Generation of new user keys and the new unit/integration tests both require access to the GATK private key, which can only be read by members of the group "gsagit".	2012-03-06 00:09:43 -05:00
Lechu	027843d791	I've simply added a "library(grid)" call at the beginning of the R script generation since R 2.14.2 doesn't seem to load the "grid" package as default. I haven't tested it on previous R versions (you may edit the R version comment to be more precise if desired), but I'm almost certain that this library call shouldn't do any harm on them. Signed-off-by: Ryan Poplin <rpoplin@broadinstitute.org>	2012-03-05 21:27:03 -05:00
Ryan Poplin	9b53250bef	Adding Unit test for Haplotype class. Used in HC's genotype given alleles mode.	2012-03-05 21:07:36 -05:00
Ryan Poplin	b37461587d	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-03-05 17:54:59 -05:00
Ryan Poplin	c6ded4d23c	Bug fix for hard clipping reads when base insertion and base deletion qualities are present in the read. Updating HaplotypeCaller integration tests to reflect all the recent changes.	2012-03-05 17:54:42 -05:00
Ryan Poplin	14a77b1e71	Getting rid of redundant methods in MathUtils. Adding unit tests for approximateLog10SumLog10 and normalizeFromLog10. Increasing the precision of the Jacobian approximation used by approximateLog10SumLog which changes the UG+HC integration tests ever so slightly.	2012-03-05 12:28:32 -05:00
Mauricio Carneiro	e9ad382e74	unifying the BQSR argument collection	2012-03-05 10:48:26 -05:00
Ryan Poplin	f879daa7d0	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-03-05 08:29:08 -05:00
Ryan Poplin	d6871967ae	Adding more unit tests and contracts to PairHMM util class. Updating HaplotypeCaller to use the new PairHMM util class. Now that the HMM result isn't dependent on the length of the haplotype there is no reason to ensure all haplotypes have the save length which simplifies the code considerably.	2012-03-05 08:28:42 -05:00
Guillermo del Angel	3b5a7c34d7	Added argument to ValidationAmplicons to only output valid sequences - useful for not having to post-filter or grep resulting files before delivering downstream	2012-03-04 10:24:29 -05:00
Mark DePristo	69611af7d3	Workaround for bug in Picard in ReadGroupProperties -- NPE caused when you call getRunDate on a read group without a date.	2012-03-02 18:53:45 -05:00
Mark DePristo	ba71b0aee4	ReadGroupProperties mk3 -- Includes sequencing date	2012-03-02 16:12:42 -05:00
Eric Banks	1e07e97b58	Optimization: create allele list just once, not for each genotype	2012-03-02 13:30:17 -05:00
Ryan Poplin	0ad7d5fbc1	Standalone common Pair HMM utility class with associated unit tests.	2012-03-01 22:41:13 -05:00
Mark DePristo	2f334a57c2	ReadGroupProperties mk2 -- Includes paired end status (T/F) -- Includes count of reads used in calculation -- Includes simple read type (2x76 for example) -- Better handling of insert size, read length when there's no data, or the data isn't paired end by emitting NA not 0	2012-03-01 18:43:53 -05:00
Mauricio Carneiro	486712bfc2	ugly RG encoding	2012-03-01 17:56:45 -05:00
Mark DePristo	aff508e091	ReadGroupProperties walker and associated infrastructure -- ReadGroupProperties: Emits a GATKReport containing read group, sample, library, platform, center, median insert size and median read length for each read group in every BAM file. -- Median tool that collects up to a given maximum number of elements and returns the median of the elements. -- Unit and integration tests for everything. -- Making name of TestProvider protected so subclasses and override name more easily	2012-03-01 15:01:11 -05:00
Mauricio Carneiro	9e95b10789	Context covariate now operates as a highly compressed bitset * All contexts with 'N' bases are now collapsed as uninformative * Context size is now represented internally as a BitSet but output as a dna string * Temporarily disabled sorted outputs because of null objects	2012-02-29 19:25:21 -05:00
Mauricio Carneiro	d379c3763a	DNA Sequence to BitSet and vice-versa conversion tools * Turns DNA sequences (for context covariates) into bit sets for maximum compression * Allows variable context size representation guaranteeing uniqueness. * Works with long precision, so it is limited to a context size of 31 bases (can be extended with BigNumber precision if necessary). * Unit Tests added	2012-02-29 19:25:20 -05:00
Eric Banks	129b5e7f6b	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-28 10:09:34 -05:00
Eric Banks	a4a279ce80	Damn you, Mark	2012-02-28 10:09:09 -05:00
Khalid Shakir	0681bea5a5	Changed DoC from PartitionType.INTERVAL to PartitionType.NONE since it doesn't have a way to gather scattered outputs. Added MultiallelicSummary to HSP eval.	2012-02-28 09:27:27 -05:00
Eric Banks	bd398e30fd	Another quick optimization	2012-02-28 09:25:35 -05:00
Eric Banks	40bdadbda5	Minor optimization as per Mark	2012-02-28 09:24:07 -05:00
Eric Banks	d7928ad669	Drat, missed one: handle null alleles being passed in.	2012-02-27 21:31:54 -05:00
Mark DePristo	24356f11b7	Merged bug fix from Stable into Unstable -- Resolved conflict Conflicts: public/java/src/org/broadinstitute/sting/gatk/datasources/reads/SAMDataSource.java	2012-02-27 17:13:17 -05:00
Mark DePristo	0b29d54937	Changed most BAMSchedule ReviewedStingExceptions to UserExceptions -- As these represent the bulk of the StingExceptions coming from BAMSchedule and are caused by simple problems like the user providing bad input tmp directories, etc.	2012-02-27 17:08:41 -05:00
Mark DePristo	f9e8e82e33	Removed unused class variable from VCFHeaderLineTranslator	2012-02-27 17:07:19 -05:00
Mark DePristo	100ddef930	Fix typo in VariantContextBuilder	2012-02-27 17:06:45 -05:00
Mark DePristo	5f7ccdcc01	Avoid calling getBasePileup when there's no pileup in NBaseCount annotation	2012-02-27 15:12:25 -05:00
Mark DePristo	729bb954e2	Throws ReviewedStingException for a bug when parent VariantContext argument is null	2012-02-27 15:09:00 -05:00
Eric Banks	998ed8fff3	Bug fix to deal with VCF records that don't have GTs. While in there, optimized a bunch of related functions (including removing a copy of the method calculateChromosomeCounts(); why did we have 2 copies? very dangerous).	2012-02-27 14:56:10 -05:00
Mark DePristo	4d9582de77	More general catching of Exceptions in interval reading to throw MalformedFile exception in all cases -- Now throws UserException no matter what happens during the reading of the intervals file.	2012-02-27 14:02:26 -05:00
Mark DePristo	9712fed7a5	Trap SAMFormatException and rethrow as MalformatedBAM exception -- Trap errors in header and rethrow -- Wrap underlying iterator in MalformatedBAMErrorReformattingIterator	2012-02-27 13:52:50 -05:00
Eric Banks	64754e7870	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-27 11:31:41 -05:00
Eric Banks	850c5d0db2	Enabling Rank Sum Tests for multi-allelics: use ref vs any alt allele.	2012-02-27 09:59:36 -05:00
Eric Banks	dfdf4f989b	Enabling Fisher Strand for multi-allelics: use the alt allele with max AC. Added minor optimization to the method in the VC.	2012-02-27 09:50:09 -05:00
Guillermo del Angel	16122bea8d	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-25 13:57:54 -05:00
Guillermo del Angel	dea35943d1	a) Bug fix in calling new functions that give indel bases and length from regular pileup in LocusIteratorByState, b) Added unit test to cover these.	2012-02-25 13:57:28 -05:00
Mark DePristo	c8a06e53c1	DoC now properly handles reference N bases + misc. additional cleanups -- DoC now by default ignores bases with reference Ns, so these are not included in the coverage calculations at any stage. -- Added option --includeRefNSites that will include them in the calculation -- Added integration tests that ensures the per base tables (and so all subsequent calculations) work with and without reference N bases included -- Reorganized command line options, tagging advanced options with @Advanced	2012-02-25 11:32:50 -05:00
Guillermo del Angel	c9a4c74f7a	a) Bug fixes for last commit related to PileupElements (unit tests are forthcoming). b) Changes needed to make pool caller work in GENOTYPE_GIVEN_ALLELES mode c) Bug fix (yet again) for UG when GENOTYPE_GIVEN_ALLELES and EMIT_ALL_SITES are on, when there's no coverage at site and when input vcf has genotypes: output vcf would still inherit genotypes from input vcf. Now, we just build vc from scratch instead of initializing from input vc. We just take location and alleles from vc	2012-02-24 10:27:59 -05:00
Mauricio Carneiro	ee9a56ad27	Fix subtle bug in the ReduceReads stash reported by Adam * The tailSet generated every time we flush the reads stash is still being affected by subsequent clears because it is just a pointer to the parent element in the original TreeSet. This is dangerous, and there is a weird condition where the clear will affects it. * Fix by creating a new set, given the tailSet instead of trying to do magic with just the pointer.	2012-02-23 18:35:25 -05:00
Mark DePristo	e0c189909f	Added support for breakpoint alleles -- See https://getsatisfaction.com/gsa/topics/support_vcf_4_1_structural_variation_breakend_alleles?utm_content=topic_link&utm_medium=email&utm_source=new_topic -- Added integrationtest to ensure that we can parse and write out breakpoint example	2012-02-23 12:14:48 -05:00
Guillermo del Angel	6866a41914	Added functionality in pileups to not only determine whether there's an insertion or deletion following the current position, but to also get the indel length and involved bases - definitely needed for extended event removal, and needed for pool caller indel functionality.	2012-02-23 09:45:47 -05:00
Eric Banks	d34f07dba0	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-22 20:41:03 -05:00
Ryan Poplin	2b6c0939ab	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-22 19:00:38 -05:00
Ryan Poplin	8695738400	Bug fix in HaplotypeCaller's GENOTYPE_GIVEN_ALLELES mode for insertions greater than length 1. The allele being genotyped was off by one base pair.	2012-02-22 19:00:04 -05:00
Christopher Hartl	2c1b14d35e	Mostly small changes to my own scala scripts: .vcf.gz compatibility for output files, smarter beagle generation, simple script to scatter-gather combine variants. Whole genome indel calling now uses the gold standard indel set.	2012-02-22 17:20:04 -05:00
Mauricio Carneiro	75783af6fc	int <-> BitSet conversion utils for MathUtils * added unit tests.	2012-02-21 14:10:36 -05:00
Guillermo del Angel	0f5674b95e	Redid fix for corner case when forming consensus with reads that start/end with insertions and that don't agree with each other in inserted bases: since I can't iterate over the elements of a HashMap because keys might change during iteration, and since I can't use ConcurrentHashMaps, the code now copies structure of (bases, number of times seen) into ArrayList, which can be addressed by element index in order to iterate on it.	2012-02-20 09:12:51 -05:00
Ryan Poplin	3d9eee4942	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-18 10:55:29 -05:00
Ryan Poplin	a8be96f63d	This caching in the BQSR seems to be too slow now that there are so many keys	2012-02-18 10:54:39 -05:00
Ryan Poplin	78718b8d6a	Adding Genotype Given Alleles mode to the HaplotypeCaller. It constructs the possible haplotypes via assembly and then injects the desired allele to be genotyped.	2012-02-18 10:31:26 -05:00
Guillermo del Angel	e724c63f2b	Reverting last commit until I learn how to effectively replicate and debug pipeline test failures, and until I also learn how to effectively remove a kep from a HashMap that's being iterated on	2012-02-17 17:18:43 -05:00
Guillermo del Angel	f2ef8d1d23	Reverting last commit until I learn how to effectively replicate and debug pipeline test failures, and until I also learn how to effectively remove a kep from a HashMap that's being iterated on	2012-02-17 17:15:53 -05:00
Guillermo del Angel	3e031a540f	Solve merge conflict	2012-02-17 10:56:03 -05:00
Guillermo del Angel	cd352f502d	Corner case bug fix: if a read starts with an insertion, when computing the consensus allele for calling the insertion was only added to the last element in the consensus key hash map. Now, an insertion that partially overlaps with several candidate alleles will have their respective count increased for all of them	2012-02-17 10:21:37 -05:00
Eric Banks	2f33c57060	No reason to restrict HaplotypeScore to bi-allelic SNPs when the plumbing for multi-allelic events is already present.	2012-02-16 13:58:00 -05:00
Guillermo del Angel	2f08846d82	Merged bug fix from Stable into Unstable	2012-02-14 21:26:25 -05:00
Guillermo del Angel	7dc6f73399	Bug fix for validation site selector: records with AC=0 in them were always being thrown out if input vcf was sites-only, even when -ignorePolymorphicStatus flag was set	2012-02-14 21:11:24 -05:00
Ryan Poplin	30085781cf	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-14 14:01:20 -05:00
Ryan Poplin	ae5b42c884	Put base insertion and base deletions in the SAMRecord as a string of quality scores instead of an array of bytes. Start of a proper genotype given alleles mode in HaplotypeCaller	2012-02-14 14:01:04 -05:00
David Roazen	85d31f80a2	Merged bug fix from Stable into Unstable	2012-02-13 16:37:11 -05:00
David Roazen	03e5184741	Fix serious engine bug that could cause reads to be dropped under certain circumstances When aggregating raw BAM file spans into shards, the IntervalSharder tries to combine file spans when it can. Unfortunately, the method that combines two BAM file spans was seriously flawed, and would produce a truncated union if the file spans overlapped in certain ways. This could cause entire regions of the BAM file containing reads within the requested intervals to be dropped. Modified GATKBAMFileSpan.union() to correct this problem, and added unit tests to verify that the correct union is produced regardless of how the file spans happen to overlap. Thanks to Khalid, who did at least as much work on this bug as I did.	2012-02-13 16:25:21 -05:00
Eric Banks	ad90af94ed	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-13 15:10:10 -05:00
Eric Banks	0920a1921e	Minor fixes to splitting multi-allelic records (as regards printing indel alleles correctly); minor code refactoring; adding integration tests to cover +/- splitting multi-allelics.	2012-02-13 15:09:53 -05:00
Eric Banks	14981bed10	Cleaning up VariantsToTable: added docs for supported fields; removed one-off hidden arguments for multi-allelics; default behavior is now to include multi-allelics in one record; added option to split multi-allelics into separate records.	2012-02-13 14:32:03 -05:00
Ryan Poplin	e9338e2c20	Context covariate needs to look in the reverse direction for negative stranded reads.	2012-02-13 13:40:41 -05:00
Ryan Poplin	41ffd08d53	On the fly base quality score recalibration now happens up front in a SAMIterator on input instead of in a lazy-loading fashion if the BQSR table is provided as an engine argument. On the fly recalibration is now completely hooked up and live.	2012-02-13 12:35:09 -05:00
Ryan Poplin	3caa1b83bb	Updating HC integration tests	2012-02-11 11:48:32 -05:00
Ryan Poplin	9b8fd4c2ff	Updating the half of the code that makes use of the recalibration information to work with the new refactoring of the bqsr. Reverting the covariate interface change in the original bqsr because the error model enum was moved to a different class and didn't make sense any more.	2012-02-11 10:57:20 -05:00
Eric Banks	f52f1f659f	Multiallelic implementation of the TDT should be a pairwise list of values as per Mark Daly. Integration tests change because the count in the header is now A instead of 1.	2012-02-10 14:15:59 -05:00
Mauricio Carneiro	1fb19a0f98	Moving the covariates and shared functionality to public so Ryan can work on the recalibration on the fly without breaking the build. Supposedly all the secret sauce is in the BQSR walker, which sits in private.	2012-02-10 11:44:01 -05:00
Eric Banks	5e18020a5f	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-10 11:08:33 -05:00
Eric Banks	f53cd3de1b	Based on Ryan's suggestion, there's a new contract for genotyping multiple alleles. Now the requester submits alleles in any arbitrary order - rankings aren't needed. If the Exact model decides that it needs to subset the alleles because too many were requested, it does so based on PL mass (in other words, I moved this code from the SNPGenotypeLikelihoodsCalculationModel to the Exact model). Now subsetting alleles is consistent.	2012-02-10 11:07:32 -05:00
Mauricio Carneiro	5af373a3a1	BQSR with indels integrated! * added support to base before deletion in the pileup * refactored covariates to operate on mismatches, insertions and deletions at the same time * all code is in private so original BQSR is still working as usual in public * outputs a molten CSV with mismatches, insertions and deletions, time to play! * barely tested, passes my very simple tests... haven't tested edge cases.	2012-02-09 18:46:45 -05:00
Eric Banks	7a937dd1eb	Several bug fixes to new genotyping strategy. Update integration tests for multi-allelic indels accordingly.	2012-02-09 16:14:22 -05:00
Eric Banks	0f728a0604	The Exact model now subsets the VC to the first N alleles when the VC contains more than the maximum number of alleles (instead of throwing it out completely as it did previously). [Perhaps the culling should be done by the UG engine? But theoretically the Exact model can be called outside of the UG and we'd still want the context subsetted.]	2012-02-09 14:02:34 -05:00
Matt Hanna	aa097a83d5	Merge branch 'master' of ssh://gsa1/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-09 11:26:48 -05:00
Matt Hanna	b57d4250bf	Documentation request by Eric. At each stage of the GATK where filtering occurs, added documentation suggesting the goal of the filtering along with examples of suggested inputs and outputs.	2012-02-09 11:24:52 -05:00
Mauricio Carneiro	d561914d4f	Revert "First implementation of GATKReportGatherer" premature push from my part. Roger is still working on the new format and we need to update the other tools to operate correctly with the new GATKReport. This reverts commit aea0de314220810c2666055dc75f04f9010436ad.	2012-02-08 23:28:55 -05:00
Eric Banks	2f800b078c	Changes to default behavior of UG: multi-allelic mode is always on; max number of alternate alleles to genotype is 3; alleles in the SNP model are ranked by their likelihood sum (Guillermo will do this for indels); SB is computed again.	2012-02-08 15:27:16 -05:00
Matt Hanna	51ac87b28c	Merge branch 'master' of ssh://gsa1/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-08 08:43:55 -05:00
Matt Hanna	5b58fe741a	Retiring Picard customizations for async I/O and cleaning up parts of the code to use common Picard utilities I recently discovered. Also embedded bug fix for issues reading sparse shards and did some cleanup based on comments during BAM reading code transition meetings.	2012-02-08 08:34:37 -05:00
Roger Zurawicki	c0c676590b	First implementation of GATKReportGatherer - Added the GATKReportGatherer - Added private methods in GATKReport to combine Tables and Reports - It is very conservative and it will only gather if the table columns, match. - At the column level it uses the (redundant) row ids to add new rows. It will throw an exception if it is overwriting data. Added the gatherer functions to CoverageByRG Also added the scatterCount parameter in the Interval Coverage script Made some more GATKReport methods public The UnitTest included shows that the merging methods work Added a getter for the PrimaryKeyName Fixed bugs that prevented the gatherer form working Working GATKReportGatherer Has only the functional to addLines The input file parser assumes that the first column is the primary key Signed-off-by: Mauricio Carneiro <carneiro@broadinstitute.org>	2012-02-07 18:14:47 -05:00
Mauricio Carneiro	e89887cd8e	laying groundwork to have insertions and deletions going through the system.	2012-02-07 18:11:53 -05:00
Mauricio Carneiro	0d3ea0401c	BQSR Parameter cleanup * get rid of 320C argument that nobody uses. * get rid of DEFAULT_READ_GROUP parameter and functionality (later to become an engine argument).	2012-02-07 14:42:11 -05:00
Eric Banks	717cd4b912	Document -L unmapped	2012-02-07 13:30:54 -05:00
Eric Banks	718da7757e	Fixes to ValidateVariants as per GS post: ref base of mixed alleles were sometimes wrong, error print out of bad ACs was throwing a RuntimeException, don't validate ACs if there are no genotypes.	2012-02-07 13:15:58 -05:00
Eric Banks	9d1a19bbaa	Multi-allelic indels were not being printed out correctly in VariantsToTable; fixed.	2012-02-06 22:49:29 -05:00
Mauricio Carneiro	5961868a7f	fixup for BQSR (HC integration tests) In the new BQSR implementation, covariates do depend on the RecalibrationArgumentCollection.	2012-02-06 22:47:27 -05:00
Mauricio Carneiro	6e6f0f10e1	BaseQualityScoreRecalibration walker (bqsr v2) first commit includes * Adding the context covariate standard in both modes (including old CountCovariates) with parameters * Updating all covariates and modules to use GATKSAMRecord throughout the code. * BQSR now processes indels in the pileup (but doesn't do anything with them yet)	2012-02-06 17:38:29 -05:00
Eric Banks	0717c79901	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-06 16:23:36 -05:00
Eric Banks	91897f5fe7	Transpose rows/cols in AF table to make it molten (so I can plot easily in R)	2012-02-06 16:23:32 -05:00
Guillermo del Angel	fb5786385c	Merged bug fix from Stable into Unstable	2012-02-06 13:22:56 -05:00
Guillermo del Angel	6ec686b877	Complement to previous commit: make sure we also don't inherit filter from input VCF when genotyping at an empty site	2012-02-06 13:19:26 -05:00
Guillermo del Angel	93ffca1e3a	Merged bug fix from Stable into Unstable	2012-02-06 11:58:58 -05:00
Guillermo del Angel	827be878b4	Bug fix when running UG in GenotypeGivenAlleles mode: if an input site to genotype had no coverage, the output VCF had AC,AF and AN inherited from input VCF, which could have nothing to do with given BAM so numbers could be non-sensical. Now new vc has clear attributes instead of attributes inherited from input VCF.	2012-02-06 11:58:13 -05:00
Eric Banks	fbbd04621d	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-06 11:53:31 -05:00
Eric Banks	edb4edc08f	Commented out unused metrics for now	2012-02-06 11:53:15 -05:00
Ryan Poplin	096c23a473	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-06 11:10:38 -05:00
Ryan Poplin	dc05b71e39	Updating Covariate interface with Mauricio to include an errorModel parameter. On the fly recalibration of base insertion and base deletion quals is live for the HaplotypeCaller	2012-02-06 11:10:24 -05:00
Guillermo del Angel	1e11408f8b	Merged bug fix from Stable into Unstable	2012-02-06 10:34:26 -05:00
Guillermo del Angel	090d87b48b	Bug fix in ValidationSiteSelector: when input vcf had genotypes and was multiallelic, the parsing of the AF/AC fields was wrong. Better logic to unify parsing of field	2012-02-06 10:33:12 -05:00
Eric Banks	9d94f310f1	Break AF histogram into max and min AFs	2012-02-06 09:01:19 -05:00
Ryan Poplin	b7ffd144e8	Cleaning up the covariate classes and removing unused code from the bqsr optimizations in 2009.	2012-02-06 08:54:42 -05:00
Eric Banks	cef550903e	Minor optimization	2012-02-06 00:48:00 -05:00
Ryan Poplin	5343f8ba67	Initial version of on-the-fly, lazy loading base quality score recalibration. It isn't completely hooked up yet but I'm committing so Mauricio and Mark can see how I envision it will fit together. Look it over and give any feedback. With the exception of the Solid specific code we are very very close to being able to remove TableRecalibrationWalker from the code base and just replace it with PrintReads -BQSR recal.csv	2012-02-05 13:09:03 -05:00
Ryan Poplin	f94d547e97	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-03 17:14:20 -05:00
Ryan Poplin	894d3340be	Active Region Traversal should use GATKSAMRecords everywhere instead of SAMRecords. misc cleanup.	2012-02-03 17:13:52 -05:00
Mauricio Carneiro	4a57add6d0	First implementation of DiagnoseTargets * calculates and interprets the coverage of a given interval track * allows to expand intervals by specified number of bases * classifies targets as CALLABLE, LOW_COVERAGE, EXCESSIVE_COVERAGE and POOR_QUALITY. * outputs text file for now (testing purposes only), soon to be VCF. * filters are overly aggressive for now.	2012-02-03 17:12:43 -05:00
Mauricio Carneiro	3dd6a1f962	Adding some generic sum and average functions to MathUtils	2012-02-03 17:12:43 -05:00
Mauricio Carneiro	e1d69e4060	make the size of a GenomeLoc int instead of long it will never be bigger than an int and it's actually useful to be an int so we can use it as parameters to array/list/hash size creation.	2012-02-03 17:12:42 -05:00
Ryan Poplin	0e44430e47	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-03 13:45:11 -05:00
Christopher Hartl	aa3638ecb3	Merge branch 'master' of ssh://chartl@ni.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-03 13:42:09 -05:00
Eric Banks	3abfbcbcf2	Generalized the TDT for multi-allelic events	2012-02-03 12:23:21 -05:00
Ryan Poplin	601e53d633	Fix when specifying preset active regions with -AR argument	2012-02-02 16:34:26 -05:00
Christopher Hartl	0111505ea9	Terrible. Swapping the paternal and sample ids.	2012-02-02 11:41:16 -05:00
Ryan Poplin	1f50f6970b	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-02 10:17:13 -05:00
Ryan Poplin	4ed06801a7	Updating HaplotypeCaller's HMM calc to use GOP as a function of the read instead of a function of the haplotype in preparation for IQSR	2012-02-02 10:17:04 -05:00
Matt Hanna	8adfc79123	Merged bug fix from Stable into Unstable	2012-02-01 16:07:41 -05:00
Matt Hanna	30b937d2af	Fix bug discovered in FGTP branch in which BlockInputStream returns -1 in cases where some data could be read, but not all the data requested by the caller.	2012-02-01 16:06:22 -05:00
Mauricio Carneiro	45da892ecc	Better exceptions to catch malformed reads * throw exceptions in LocusIteratorByState when hitting reads starting or ending with deletions	2012-02-01 11:56:19 -05:00
Christopher Hartl	810996cfca	Introducing: VariantsToPed, the world's most annoying walker! And also a busted QScript to run it that I need Khalid's help debugging ( frownie face ). Note that VariantsToPed and PlinkSeq generate the same binary file (up to strand flips...thanks PlinkSeq), so I know it's working properly. Hooray!	2012-02-01 10:39:03 -05:00
Christopher Hartl	25d943f706	Merge branch 'master' of ssh://chartl@ni.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-02-01 10:32:11 -05:00
Ryan Poplin	056b24ccd6	Resolving merge conflicts with LocusIteratorByState	2012-01-31 16:13:32 -05:00
Ryan Poplin	febc634557	Changing PileupElement's isSoftClipped to isNextToSoftClip since soft clipped bases aren't actually added to pileups, oops. Removing the intrinsic clustered variants filter from the HaplotypeCaller	2012-01-31 16:06:14 -05:00
Matt Hanna	7f70612beb	Merge branch 'master' of ssh://gsa1/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-01-31 11:59:25 -05:00
Matt Hanna	a630db1703	Oops...HierarchicalMicroScheduler was transforming any exception from the walker level into a ReviewedStingException. Thanks to Ryan for pointing this out.	2012-01-31 11:58:21 -05:00
Christopher Hartl	faba3dd530	Merge branch 'master' of ssh://chartl@ni.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-01-31 10:25:29 -05:00
Mauricio Carneiro	17dbe9a95d	A few cleanups in the LocusIteratorByState * No more N's in the extended event pileups * Only add to the pileup MQ0 counter if the read actually goes into the pileup	2012-01-31 09:40:51 -05:00
Ryan Poplin	f9162ea705	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-01-30 19:45:19 -05:00
Ryan Poplin	abb91cf26b	Increasing the size of the active regions that are produced by the active probability integrator, more context is needed to call more complex events	2012-01-30 15:36:12 -05:00
Mauricio Carneiro	d5d4fa8a88	Fixed discordance bug reported by Brad Chapman discordance now reports discordance between genotypes as well (just like concordance)	2012-01-30 09:50:45 -05:00
Mark DePristo	3164c8dee5	S3 upload now directly creates the XML report in memory and puts that in S3 -- This is a partial fix for the problem with uploading S3 logs reported by Mauricio. There the problem is that the java.io.tmpdir is not accessible (network just hangs). Because of that the s3 upload fails because the underlying system uses tmpdir for caching, etc. As far as I can tell there's no way around this bug -- you cannot overload the java.io.tmpdir programmatically and even if I could what value would we use? The only solution seems to me is to detect that tmpdir is hanging (how?!) and fail with a meaningful error.	2012-01-29 15:14:58 -05:00
Menachem Fromer	0e17cbbce9	Merged bug fix from Stable into Unstable	2012-01-27 16:03:16 -05:00
Menachem Fromer	a9671b73ca	Fix to permit proper handling of mapping qualities between 128 to 255 (which get converted to byte values of -128 to -1)	2012-01-27 16:01:30 -05:00
Ryan Poplin	f7ac1f4a69	Merge branch 'master' of ssh://nickel.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-01-27 15:12:55 -05:00
Ryan Poplin	fc08235ff3	Bug fix in active region traversal, locusView.getNext() skips over pileups with zero coverage but still need to count them in the active probability integrator	2012-01-27 15:12:37 -05:00
Mark DePristo	0f2e8400b5	Merge branch 'master' of ssh://gsa1/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-01-27 10:12:50 -05:00
Mauricio Carneiro	ec9920b04f	Updating the SAM TAG for Original Alignment Start to "OP" per Mark's recommendation to reuse the Indel Realigner tag that made it to the SAM spec. The Alignment end tag is still "OE" as there is no official tag to reuse.	2012-01-27 08:51:39 -05:00
Mark DePristo	13d1626f51	Minor improvements in ref QC walker. Unfortunately this doesn't actually catch Chris's error	2012-01-27 08:24:22 -05:00

... 7 8 9 10 11 ...

2168 Commits (6f8e7692d4e67d5cfb2d2f9e33549ff4fbf2e5f8)