Releases: pdimens/harpy
Release list
4.2
New
QC
- replace multiqc report with native harpy report
Align
- BWAMEM2 has been replaced with minibwa. Long live BWA!
- it's much faster, and takes much less time to index a reference
- coverage depth added to aggregate report for processed alignments
- minimap2 added back in (with extra perks) for long-read compatability
- called with
harpy align minimap
- called with
Report
- harpy reports can be converted to less-nice but functional standalone HTML files
- this feature is accessed using
harpy report static - to accomodate this,
harpy report(live report website) is nowharpy report live
- this feature is accessed using
Changes
Align
-d(molecule distance threshold) has its default restored to 50kb since this value is used exclusively for reporting and does not alter the data
misc
- removed FASTA format validation because it can be dreadfully slow with existing tools
Report
- harpy reports can be converted to less-nice but functional standalone HTML files
- this feature is accessed using
harpy report static - to accomodate this,
harpy report(live report website) is nowharpy report live
- this feature is accessed using
Fixes
reports
- tables now render properly in VScode/Jupyter contexts
- xeus-python began to hang for unknown reasons and has been replaced with
ipykernel
misc
- constrain CASAVA regex in FASTQ file validation so it doesn't trigger false positives when new CASAVA appears in unexpected places
- [internal] notebooks no longer a submodule/subdirectory of
harpy.report - utility
check_fastq.pyno longer employs globals, instead uses a sensible class system - add multithreading to pre-workflow VCF and XAM file validation and parsing
- FASTA validation no longer checks the entire file for formatting b/c it was taking much too long on bigger genomes
- still considering alternatives, but is disabled for now
- [internal-ish] the bwa, strobealign, and minimap2 workflows are nearly identical except for the reference preprocessing and alignment, so to minimize redundancy and duplication, those workflows have a single
align.smkthat imports a second snakefilealign_{aligner}.smkthat handles just the preprocessing and direct alignment for those aligners, then hands off toalign.smkfor all the downstream things (dedup, sorting, reports, etc)
Documentation
- the pages for bwa, strobealign, and minimap have been consolidated into a single page bc they are nearly identical
What's Changed
- integrate qc report natively by @pdimens in #289
- Standalone reports by @pdimens in #291
- Bump the dependencies group with 2 updates by @dependabot[bot] in #293
- aligners by @pdimens in #294
Full Changelog: 4.1.5...4.2
4.1.5
Fixes
- #285 : remove 2nd pie chart in linked-read alignment reports
- #287 when using
--vcf-samples, filter input bam files with VCF sample list
PRs
Full Changelog: 4.1.4...4.1.5
4.1.4
Fixes
Snakefiles
- move output notebook within shell to avoid xpython stdout printing error breaking notebooks
Full Changelog: 4.1.3...4.1.4
4.1.3
Fixes
Reports
- preprocess GIH uses log-scaled binning to reduce data significantly and prevent altair crashes
- align aggregate linked read Total vs. plot uses reversed tubro palette and a dropdown selector instead of radio
- samtools stats stats-boxes sets more realistic cutoffs for %mapped, %properly paired, %optical dupes
- coerce bins in bcftools report into numeric type
- force columns pertaining to contigs into String types
- fixes altair compatability for nominal data
- fixes table parsing errors that arise when contigs are both named as a number and as strings (e.g. a contig named '1' and another named 'X')
Full Changelog: 4.1.2...4.1.3
4.1.2
Fixes
Align
- report forces
contigcolumns to be Strings, avoiding type-errors where contig names are a mix of numbers and alphanumeric
Full Changelog: 4.1.1...4.1.2
4.1.1
Fixes
QC
- adapter trimming was being incorrectly skipped in all cases
- set max threads for fastp jobs to 4
Align
- both
strobealignandbwaworkflows have the rules shifted for better memory management- specifically, the pipes between align->fixmates->markdups have been broken for speed reasons (it was very slow)
- the new rules create temporary (uncompressed BAM) files at the choke-point steps (i.e. collate, sort, markdups) to make better use of time and computational resources
samtools sortis now given more threads and RAM per thread, with the memory per thread decreasing on failed attempts- 3 total attempts. Initially 3GB RAM per thread (x4 threads), drops by half each attempt (e.g. 12GB, 6GB, 3GB total)
samtools statsproperly ignores duplicates on processed alignmentsharpy-utils optical-distancenow properly falls back to100harpy-utils molecule-coverageis much less RAM hungryharpy-utils bx-stats-samis much less RAM hungry
Assembly/Metassembly
- BUSCO now scrubs the orthodb version from the expected output
.txtfile to prevent orthodb version updates breaking the workflow
What's Changed
- Bump the dependencies group with 2 updates by @dependabot[bot] in #281
- tolerate intermediates for speedup by @pdimens in #282
Full Changelog: 4.1...4.1.1
4.1
Fixes
harpy reporterror when not used in a git-enable repository- embedded image viewer has cleaned up HTML
- no more double-printing worklow name when using
harpy resume --unlinkedworks correctly inphase bam
Features
Impute
- the
--bufferoption includes fractional scaling when <= 1- e.g.
--strategy window:100000 -b 0.1equivalent to--strategy window:100000 -b 10000 - e.g.
--strategy contig1:1-100000 -b 0.1equivalent to--strategy contig1:1-100000 -b 10000 - only applies to window and region strategies
- e.g.
Preprocess GIH
- adds
adapters.fastaoutput to be used as input intoharpy qc
Changes
Reports
- tables have "Export CSV" replaced with "⤓ Download CSV" to be more obvious
Breaking
None
Full Changelog: 4.0...4.1
4.0
This release has been in development for 6 months (sorry for the delay!) and that's because it has so many changes and new features. The docs still need to be updated, but please review the changes in 4.0 if you're upgrading from previous versions.
New
new commands
diagnosenow has 2 subcommands:stall: same as previousdiagnosebehavior, where it runs snakemake with--dry-run --debug-dagrule: attempt to directly run the failing rule of a workflow as identified in the snakemake log, will attempt to run snakemake to generate missing inputs if necessary
phase bamadded for more fine-tuned and configurable alignment phasingreportadded to render new ipython reports as a MySTmd websiteharpy viewnow has the aliashv- when erroring, harpy will create hidden file
.harpyerrorwith the name of the directory associated with the last error - calling hv config/log/error/profile/snakefile without arguments is allowed and will default to the directory in
.harpyerror - this now streamlines troubleshooting to: harpy workflow terminates with error ->
hv logorhv config
- when erroring, harpy will create hidden file
- all user-accessible harpy utility scripts now live under
harpy-utils, which has its own CLI and is a consistent access point than remembering the names of specific utilities you may want to use- the new Go utilities are internal (for now) and not exposed by
harpy-utils
- the new Go utilities are internal (for now) and not exposed by
new options
-
resumehas new--directoption to call Snakemake directly without harpy intervention -
hidden common option
--cleanwith the optionsw,s, and/orl, to remove theworkflow/,.snakemake/, and/orlogs/directories in the output- this option is hidden because it's meant more for development
- options provided as sequential letters (e.g.
ws,sl,lw, etc.)
-
the
-T/--notempoption in snakemake is exposed in harpy commands as--no-tempto simplify using it -
all common workflow options have a capital letter short-name
short long -H--hpc-C--container-Q--quiet-S--snakemake-R--skip-reports-N--setup
miscellaneous new
- output log of checks and validations printed to console so Harpy is transparent about any observed delays before kicking off Snakemake
- disabled when
--quiet> 0
- disabled when
- progress bar has a new column to show a count of the active jobs!
- now featuring Go code!
- some utilities replaced with newly-written Go programs
- future versions will have bottlenecking utility scripts replaced with Go counterparts
- this will be a slow process as I learn more Go
- time elapsed column in progress bar pauses when there are no active jobs for that rule (better reflecting the actual time elapsed)
- added "workflow setup complete" text when using
--setup - added "all stuff is there" equivalent text when snakemake reports there is nothing to do
Deprecations
convert(replaced byDjinnsoftware)downsample(replaced byDjinnsoftware)simulate(replaced byMimickandVISOR-HACkssoftwares)
Renamed
config.yamlis nowprofile.yamlto minimize confusion and make it aligned with latest Snakemake versionview snakeparamshas been updated toview profileto match this
demultiplexis nowpreprocessto better reflect what the commands do--setup-onlyreplaced with the more succinct--setup- short-name for
--threadsis now-@(nod to htslib) --output-dirreplaced with--output- short-name for
--outputis now-O
- short-name for
Workflows
workflow configs
- the
workflow.yamlfiles now all have a standard/consistent format with three main sections whose names are capitalized (whereas all the rest are lowercase):Workflow: with common information (name, linkedread info, report skip/contigs, harpy-specific snakemake things)Parameters: the run configurations resulting from command-line arguments/optionsInputs: the input files- this means previous
workflow.yamlfiles are incompatible with this and future versions
- all
workflow.yamlkeys use hyphens instead of underscores to reduce keystrokes- e.g.
min_len=>min-len
- e.g.
- hpc.yaml profile parameters now merged into
profile.yamlso all the profile information exists in a single file (internal change) - workflow yaml files now include a
VERSIONvariable that syncs with the Harpy version used- populates the
containerversion tags within snakefile rules - this makes container versions reliable by default, but entirely hackable to manually dissociate harpy version and container version
- populates the
harpy align
- align workflows now use
mosdepthto calculate depth - alignment stats are generated from raw alignment records, better reflecting the raw alignment performance of data
- aggregate report correctly reports %linked reads rather than molecules
harpy diagnose
diagnoseis nowdiagnose stallto accomodate distinction from newdiagnose rule
harpy phase
phasehas been renamedphase snpto accommodate a disctinction from the newphase bamworkflow
harpy qc
--min-lengthand--max-lengthconsolidated into--length min,max- minimum mapping quality has been consistenlty named
min-map-qualityin allworkflow.yamlfiles
harpy sv
sv naibrno longer phases input alignment files, use the newphase bammodule for that- 4 SV reports consolidated into 1
Reports
Reports have been completely rewritten (for the third time), moving away from R/Quarto to Python/Jupyter. This change was necessary to achieve specific quality-of-life improvements that were not possible with the current setup:
- report generation is significantly faster because Quarto isn't trying to render each as a standalone HTML file
- IPython notebooks store the results within themselves, which can be accessed whenever via JupyterLab/VScode/etc
- all harpy-generated (non-MultiQC) reports can be bound and built into a singular report website using
harpy reportwith navigation between pages, searching, etc. - plots are now generated with Altair (Vega/Vegalite) plotting library, which is very fast, interactive, and reactive, with a more permissible license than Highcharts, which use a restrictive software license
Internal
- significant rewrite of the
Workflowclass and how it expects workflow, parameter, and input delcarations - printing functions consolidated into
HarpyPrintclass - workflow error reporting rewritten:
- now includes harpy version
- not relying on
whileloops of snakemake output (less likely for infinite hang after error) - visual overhaul to make it less overwhelming
- swapped order of validations/checks
- CLI input validations are for fast basic checks (e.g. naming conventions, presence/absence)
- harpy validations are for more involved checks (formatting, consistency, inter-parameter checks)
- Snakemake process monitoring saw a significant rewrite
- the new internals are easier to develop and should hopefully have more consistent exiting behavior
resumelogic reorganized updated to match new workflow configuration design- the
--workflow-profilepart of the snakemake command (when using hpc) has been moved toprofile.yamlto further reduce the length of the snakemake call - grouped and single-sample leviathan variant calling now use a single consolidated snakefile
- grouped and single-sample naibr variant calling now use a single consolidated snakefile
- workflow info printing handled differently to be more flexible
Fixes
- more accurate optical duplication detection by setting the distance parameter using
harpy-utils optical-dist, which has a more comprehensive lookup of instrument codes - removed redundant validations between CLI checks and harpy checks
- wording improvements for errors and doc text
- [hopefully] no more double-printing of Snakemake errors
VScode + xeus-python bug
Using the xeus-python kernel for the reports includes an odd edge case that can be easily avoided.
TL;DR
The root cause is essentially VS Code trying to upsell its Native REPL to users it thinks might benefit from it, without properly scoping that behavior away from already-running kernel sessions.
Fix
Change VScode python.terminal.shellIntegration.enabled setting to false:
"python.terminal.shellIntegration.enabled": falseTechnical Details
VS Code's Python extension has a feature called the Native REPL, which is an integrated interactive environment that it tries to promote when it detects certain kernel types. When you use xeus-python, VS Code identifies it as a non-standard kernel (since it's not the default IPython/ipykernel) and injects that "Ctrl+click to launch VS Code Native REPL" message as a kind of prompt/advertisement at the top of the notebook or interactive window output.
This injection happens at the extension level, not from xeus-python itself — the message is being inserted by the Pylance or Python VS Code extension into the output stream before your kernel's actual output. Chaos and sadness ensues.
3.2
New
--chooseoption forharpy view log
Fixed
- removed custom snakemake logging and updated
harpy view logto point to.snakemake/log(fixes #269) - all commands now have
--helpoption (fixes #268)
Internal
- conda environments now have a
HarpyEnvsclass that unified conda/pixi/container things - harpy containers now use smaller pixi environments, one per corresponding conda env
- resulting in more containers that are significantly smaller, but still version-tagged (e.g.
pdimens/harpy:qc_3.2)
- resulting in more containers that are significantly smaller, but still version-tagged (e.g.
Breaking Changes
None
What's Changed
Full Changelog: 3.1...3.2
3.1
deprecations
- harpy convert
- harpy downsample
- harpy simulate linkedreads
changes
- simplified the rich-click theming
- bwa-mem2 replaces bwa in align bwa
- impute needs a minimum of 5 biallelic snps per contig
- alignment during metassembly no longer outputs unmapped reads or alignments with mapq < 10
- added hidden
--forceoption to metassembly to force athena to run even if fq/bam don't pass its internal qc - options error borders are now yellow, making it consistent with other errors
fixes
- impute workflow has explicit output plot filename declarations to catch errors better
- github action that builds the container
In progress but not ready yet
- replace manually hacked progressbars with executor plugin