gatk-3.8

Commit Graph

Author	SHA1	Message	Date
Mark DePristo	d6e6b30caf	Initial implementation of GSA-515: Nanoscheduler – Write general NanoScheduler framework in utils.threading. Test with reading via iterator from list of integers, map is int * 2, reduce is sum. Should be efficiency using resources to do sum of 2 * (sum(1 - X)). Done! CPU parallelism is nano threads. Pfor across read / map / reduce. Use work queue to implement. Create general read map reduce framework in utils. Test parallelism independently before hooking up to Locus iterator Represent explicitly the dependency graph. Scheduler should choose the work units that are ready for computation, that are marked as "completing a computation", and then finally that maximize the number of sequent available work units. May be worth measuring expected cost for read read / map / reduce unit and use it to balance the compute As input is single threaded just need one thread to populate inputs, which runs as fast as possible on parallel pushing data to fixed size queue. Each push creates map job and links to upcoming reduce job. Note that there's at most one thread for IO tasks, and all of the threads can contribute to CPU tasks	2012-08-24 14:07:44 -04:00
Mark DePristo	b3fd74f0c4	HaplotypeCaller forbids BAQ	2012-08-24 13:25:05 -04:00
Mark DePristo	c689d6dcac	GATKPerformanceOverTime is finalized -- Update BQSR to run v1 and v2. Use new single read group extracted BAM -- Bug fixes	2012-08-24 09:20:32 -04:00
Mark DePristo	8371362f3c	Refactor and cleanup GATKPerformanceOverTime -- Use single read group BAM file for BQSR -- Implement terrible (but clever) hack to support BQSR v1 and v2 in a single Scala class.	2012-08-23 21:11:15 -04:00
Mark DePristo	b6cc615890	MySQLdb required to run analyzeRunReports, despite my best efforts	2012-08-23 21:08:32 -04:00
Mark DePristo	1999b95754	Work around for GSA-513: ClassCastException in VariantEval	2012-08-23 18:14:49 -04:00
Christopher Hartl	f1166d6d00	Spotted a potential bug where sample IDs passed in from the meta data were only checked against the sample IDs in the VCF header if the input file happened to be a meta data file rather than a fam file. Added a check for fam files as well, and added an integration test to cover each case.	2012-08-23 11:43:19 -07:00
Mark DePristo	2ae5ec5596	Update for GSA-506: Add nt and efficiency information to GATKRunReport -- Python log upload now includes efficiency information in GATKLogs DB	2012-08-23 12:53:22 -04:00
Mark DePristo	857b11b26f	Done with GSA-506: Add nt and efficiency information to GATKRunReport -- GATKRunReports contain itemized information about the numThreads used to execute the GATK, as well as the efficiency of the use of those threads to get real work done, including time spent running, waiting, blocking, and waiting for IO -- See https://jira.broadinstitute.org/browse/GSA-506 for more details	2012-08-23 09:59:53 -04:00
Mark DePristo	0b735884db	Cleanup code in VariantContext	2012-08-23 09:59:53 -04:00
Mark DePristo	d973863039	GATKPerformanceOverTime includes longer running tests for select variants and variant eval	2012-08-23 09:59:53 -04:00
Eric Banks	e5df91aa23	Looks like the @WalkerName annotation doesn't work with the GATK docs, so I'm renaming the walkers.	2012-08-22 20:17:39 -04:00
Mark DePristo	95a1337285	Merge branch 'threadMonitors' Conflicts: private/scala/qscript/org/broadinstitute/sting/queue/qscripts/performance/GATKPerformanceOverTime.scala	2012-08-22 16:54:47 -04:00
Mark DePristo	63af0cbcba	Cleanup GATK efficiency monitor classes -- Invert logic in GATKArgumentCollection to disable monitoring, not enable. That means monitoring is on by default -- Fix testing error in unit tests -- Rename variables in ThreadAllocation to be clearer	2012-08-22 16:48:02 -04:00
Mark DePristo	1d47d2b573	Fix GATKPerformanceOverTime for BQSR file path error	2012-08-22 16:48:02 -04:00
Mark DePristo	e1293f0ef2	GSA-507: Thread monitoring refactored so it can work without a thread factory -- Old version StateMonitoringThreadFactory refactored into base class ThreadEfficiencyMonitor and subclass EfficiencyMonitoringThreadFactory. -- Base class is used by LinearMicroScheduler to monitor performance of GATK in single threaded mode -- MicroScheduler now handles management of the efficiency monitor. Includes master thread in monitor, meaning that reduce is now included for both schedulers	2012-08-22 16:48:01 -04:00
Mark DePristo	f876c51277	Separately track time spent doing user and system CPU work -- Allows us to ID (by proxy) time spent doing IO -- Refactor StateMonitoryingThreadFactory to use it's own enum, not Thread.State -- Reliable unit tests across mac and unix	2012-08-22 16:48:01 -04:00
Mark DePristo	18060f237b	Add thread efficiency monitoring to GATK HMS -- See https://jira.broadinstitute.org/browse/GSA-502 -- New command line argument -mt enables thread monitoring -- If enabled, HMS uses StateMonitoringThreadFactory to create monitored threads, and prints out an efficiency report when HMS exits, telling the user information like: for BQSR – known to be inefficient locking INFO 17:10:33,195 StateMonitoringThreadFactory - Number of activeThreads used: 8 INFO 17:10:33,196 StateMonitoringThreadFactory - Total runtime 90.3 m INFO 17:10:33,196 StateMonitoringThreadFactory - Fraction of time spent blocked is 0.72 ( 64.8 m) INFO 17:10:33,197 StateMonitoringThreadFactory - Fraction of time spent running is 0.26 ( 23.7 m) INFO 17:10:33,197 StateMonitoringThreadFactory - Fraction of time spent waiting is 0.02 ( 112.8 s) INFO 17:10:33,197 StateMonitoringThreadFactory - Efficiency of multi-threading: 26.19% of time spent doing productive work for CountLoci INFO 17:06:12,777 StateMonitoringThreadFactory - Number of activeThreads used: 8 INFO 17:06:12,777 StateMonitoringThreadFactory - Total runtime 43.5 m INFO 17:06:12,778 StateMonitoringThreadFactory - Fraction of time spent blocked is 0.00 ( 4.2 s) INFO 17:06:12,778 StateMonitoringThreadFactory - Fraction of time spent running is 1.00 ( 43.3 m) INFO 17:06:12,779 StateMonitoringThreadFactory - Fraction of time spent waiting is 0.00 ( 6.0 s) INFO 17:06:12,779 StateMonitoringThreadFactory - Efficiency of multi-threading: 99.61% of time spent doing productive work	2012-08-22 16:48:01 -04:00
Mark DePristo	27842ba448	run_performance_tests use bsub and gsa -- Confirmed that running on gsa queue is fine with sufficient iterations (3)	2012-08-22 16:48:01 -04:00
Mark DePristo	d7a6cd99cd	Expand intervals processed for many GATKPerformanceOverTime commands -- For the high NT tests the total runtime may be too short to really assess nt efficiency vs. start up costs. Reworked underlying test data and intervals so that most tests run in 10-20 hrs for -nt 1.	2012-08-22 16:48:01 -04:00
Guillermo del Angel	1aa856e0e3	Merge branch 'master' of ssh://gsa4.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-08-22 15:53:47 -04:00
Guillermo del Angel	e29469eeeb	Forgot to update 2 integration test md5's (in this cases, changes are legit because of the code revamp of AD, it's simpler if AD is not output when a site is not variant, as genotype DP conveys the same information)	2012-08-22 15:53:33 -04:00
Menachem Fromer	b1b9c0b132	Merge branch 'master' of ssh://gsa2.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-08-22 15:26:39 -04:00
Ryan Poplin	fe3069b278	Merged bug fix from Stable into Unstable	2012-08-22 14:40:34 -04:00
Ryan Poplin	e5cfdb4811	Bug fix for popular _Duplicate allele added to VariantContext_ error reported on the forum. It seems to be due to lower case bases in the reference being treated as reference mismatches. We would try to turn these mismatches into SNP events, for example c/C. We now uppercase the result from IndexedFastaSequenceFile.getSubsequenceAt()	2012-08-22 14:39:35 -04:00
Ryan Poplin	63213e8eb5	Expanding the HaplotypeCaller integration tests to cover a wider range of data	2012-08-22 14:18:44 -04:00
Eric Banks	944e1c299d	Docs for --keepOriginalAC were wrong in SelectVariants	2012-08-22 13:07:13 -04:00
Eric Banks	2409aa9bfd	Merge branch 'master' of ssh://gsa2.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-08-22 12:54:43 -04:00
Eric Banks	94540ccc27	Using the simple VCBuilder constructor and then subsequently trying to modify attributes was throwing a NPE. This is easily solved (without a performance hit) by initializing the attributes map to an immutable Collections.emptyMap(). Added unit test to cover this case.	2012-08-22 12:54:29 -04:00
Guillermo del Angel	901f47d8af	Final step (for now) in VA refactoring: update MD5's because, a) since it's not guaranteed that we'll iterate through reads/pileups in the same order, the rank sum dithering will change annotations, b) FS uses new generic threshold to distinguish uninformative reads (it used to use ad-hoc thresholds), c) AD definition changed and throws away uninformative reads, d) shortened general ploidy integration tests for quicker debugging. May have missed some MD5's in the update so there may be lingering test failures still	2012-08-22 11:38:51 -04:00
Guillermo del Angel	7df0abf49b	Merge branch 'master' of ssh://gsa4.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-08-22 11:36:41 -04:00
Eric Banks	9e76e8aa0b	Just noticed that the efficient conversion to uppercase method is redundant since it's already implemented efficiently in Picard; let's just have a single implementation.	2012-08-22 11:26:08 -04:00
Christopher Hartl	20601f034e	Updating the checkType() function to include the new StructuralIndel variant type. Fixes outstanding broken integration test.	2012-08-22 07:33:10 -07:00
Eric Banks	c7ce3e1cf5	Merged bug fix from Stable into Unstable	2012-08-22 00:24:40 -04:00
Eric Banks	03017855e4	WTF - why is support for whole-read insertions all messed up in LIBS? I've pushed a temporary patch for now (the right solution should certainly not be implemented in stable; LIBS needs to be better thought out). Added another unit test.	2012-08-22 00:24:01 -04:00
Mark DePristo	1acf18aa25	run_performance_tests use bsub and gsa -- Confirmed that running on gsa queue is fine with sufficient iterations (3)	2012-08-21 16:26:12 -04:00
Mark DePristo	cb9ba4f660	Expand intervals processed for many GATKPerformanceOverTime commands -- For the high NT tests the total runtime may be too short to really assess nt efficiency vs. start up costs. Reworked underlying test data and intervals so that most tests run in 10-20 hrs for -nt 1.	2012-08-21 16:25:31 -04:00
Mark DePristo	1d707e7b31	Linear and Quadratic fits for GATKPerformanceOverTime.R	2012-08-21 14:44:18 -04:00
Mark DePristo	e43ff31eab	GATKPerformance over time lives (mark II) -- Now uses new tagging capabilities so that 2.x runs will tag their logs as GATKPerformanceOverTime -- Update bamboo runs, reverting back to gsa4 (it's slower but the results are less variable -- you were right david!).	2012-08-21 14:44:18 -04:00
Mark DePristo	df36b4384c	Update analyzeRunReports.py to handle new GATK tag argument	2012-08-21 14:44:18 -04:00
Mark DePristo	6ce8016ae7	GSA-491: Add hidden tag to GATK that propagates to the GATK logs	2012-08-21 14:44:18 -04:00
Mark DePristo	9eec33ec3b	Complete GSA-497: Let Queue write out runInfo on the fly, after each job group finishes running -- Queue will incrementally now write out its jobReport.txt file whenever jobs finish running (FAIL or DONE) -- This makes it far easier to track what's going on, or to analyze incrementally performance results coming out of Queue -- Generally cleaned up the QJobsReporting code, creating a new clean class QJobsReporter that holds all of the information on what to do log and where to put into, which was previously scattered in QCommandLine and QJobReport	2012-08-21 14:44:18 -04:00
Guillermo del Angel	6a8cf1c84a	Enable and adapt HaplotypeScore and MappingQualityZero as active region annotations now that we have per-read likelihoods passed in to annotations	2012-08-21 14:35:40 -04:00
Guillermo del Angel	d0644b3565	Merge branch 'master' of ssh://gsa4.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-08-21 10:35:23 -04:00
Ryan Poplin	94e7f677ad	Merge branch 'master' of ssh://gsa2.broadinstitute.org/humgen/gsa-scr1/gsa-engineering/git/unstable	2012-08-21 10:21:47 -04:00
Guillermo del Angel	418ace463a	More merge conflict resolution	2012-08-21 10:15:52 -04:00
Ryan Poplin	10961db3ce	Another round of FindBugs fixes. Object returns its internal reference to an externally mutable array. Very dangerous.	2012-08-21 09:35:55 -04:00
Ryan Poplin	605acaae9c	Another round of FindBugs fixes. Object internally stores a reference to an externally mutable array. Very dangerous.	2012-08-21 09:33:58 -04:00
Ryan Poplin	55b7949d68	Another round of FindBugs fixes. Comparator doesn't implement Serializable.	2012-08-21 09:20:55 -04:00
Christopher Hartl	ba8622ff0d	number of stashed changes are lurking in here. In order of importance: - Fix for M_Trieb's error report on the forum, and addition of integration tests to cover the walker. - Addition of StructuralIndel as a class of variation within the VariantContext. These are for variants with a full alt allele that's >150bp in length. - Adaptation of the MVLikelihoodRatio to work for a set of trios (takes the max over the trios of the MVLR) - InsertSizeDistribution changed to use the new gatk report output (it was previously broken) - RetrogeneDiscovery changed to be compatible with the new gatk report - A maxIndelSize argument added to SelectVariants - ByTranscriptEvaluator rewritten for cleanliness - VariantRecalibrator modified to not exclude structural indels from recalibration if the mode is INDEL - Documentation added to DepthOfCoverageIntegrationTest (no, don't yell at chartl ;_; ) Also sorry for the long commit history behind this that is the result of fixing merge conflicts. Because this also fixes a conflict (from git stash apply), for some reason I can't rebase all of them away. I'm pretty sure some of the commit notes say "this note isn't important because I'm going to rebase it anyway".	2012-08-21 07:08:58 -04:00

1 2 3 4 5 ...

10361 Commits (d6e6b30caf15d2f7f64fcc1f2b710b458507f7be) All Branches Search

10361 Commits (d6e6b30caf15d2f7f64fcc1f2b710b458507f7be)

All Branches