Correctness and performance pass: mapping positions, EM convergence, Gibbs prior, SAM, GFF3, CLI - #1231
Open
BenjaminDEMAILLE wants to merge 13 commits into
Open
Correctness and performance pass: mapping positions, EM convergence, Gibbs prior, SAM, GFF3, CLI#1231BenjaminDEMAILLE wants to merge 13 commits into
BenjaminDEMAILLE wants to merge 13 commits into
Conversation
BenjaminDEMAILLE
force-pushed
the
claude/perf-and-bugfixes
branch
from
October 1, 2026 06:26
4ae22ba to
907e832
Compare
BenjaminDEMAILLE
marked this pull request as ready for review
October 1, 2026 06:27
Positions were taken from the first/last seed of the MEM chain, so an unmatched read prefix/suffix (e.g. a mismatch in the edge k-mer) shortened fragments and shifted coordinates. MappingCandidate now records the read length and exposes proj_start()/proj_end() (chain start minus unmatched prefix, chain end plus unmatched suffix), used for fragment lengths (FLD), library-format and dovetail detection, SAM POS and bias 5' coordinates. Also fixes the off-by-one 5' position for reverse single-end reads and orphans (last base is proj_end() - 1). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A pair recovered from an orphan anchor now gets the fragment's leftmost base as ref_pos, the partner's real leftmost position (a reverse anchor ends the fragment, so the partner starts it) and fragment 5'/3' ends for positional bias, as for a concordant pair. recover_mate clamps its search window to the transcript (no negative window start when the read overhangs) and reports the fragment length relative to the projected anchor start. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
For a paired fragment ref_pos is the fragment's leftmost base whichever mate is forward, so the forward 5' context is ref_pos and the reverse 5' context is the fragment's last base. The old code branched on mate 1's orientation and mis-attributed contexts for ISR-type libraries where mate 1 is the reverse read. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Fail when a reference's FASTA length differs from its length in the index (e.g. duplicate reference names) instead of writing a corrupt store. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max_rel_diff now filters on alpha_out (as salmon does) and returns -inf when no transcript clears the cutoff, which counts as converged. Before, an input with no fragments ran EM/VBEM for max_iter iterations. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rNucleotidePrior The Gibbs prior kept max(1, vbPrior) when switched to per-nucleotide mode, so it no longer matched the VBEM point estimate (vbPrior * max(1, effLen)), and plain EM runs switched to a per-nucleotide 1e-3 prior. The per-nucleotide prior now applies only with VBEM and uses vbPrior itself. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ambig_info.tsv counts were summed in u32 after truncating the u64 class counts, which can overflow on deep libraries. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A read ending exactly at the transcript start or starting exactly at its end produced a 0M op (rejected by htslib/samtools); it is now written as all soft-clip. Adds unit tests for overhang_cigar. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A zero-length fragment has no end base; starting the convolution at length 0 indexed position -1 and panicked. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Ensembl GFF3 puts gene_id only on gene features and links transcripts to them with Parent=<gene ID>. Transcripts without gene_id are now resolved through their Parent after the pass, falling back to the bare identifier of gene:ENSG... IDs when the gene feature is missing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Reject non-positive or non-finite --fldMean/--fldSD, GC bin counts outside 1..=101, --thinningFactor 0 and out-of-range scoring flags at parse time. --mp accepts salmon's negative convention (--mp -4); its magnitude is used. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each Gibbs chain keeps one probs/draws scratch pair for all its rounds instead of allocating a draws vector per multi-transcript class per round (and a probs vector per round). Also adds a test that parallel and sequential EM/VBEM agree. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
add_val accumulates the fragment's total and length-weighted mass locally and publishes them with one atomic update each, instead of a CAS loop on the shared sum/tot_mass per kernel entry, avoiding cache-line contention between workers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
BenjaminDEMAILLE
force-pushed
the
claude/perf-and-bugfixes
branch
from
October 1, 2026 09:15
907e832 to
5053d6f
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rebased onto current
master. Changes that master already covers (decoy block located by name, deterministic eq-class order, replicate scaling, parallel digamma and balanced EM shards, allocation-free alignment cache, posterior reuse) were dropped. One commit per change.Correctness
approxReadStartPos), not from the first/last seed. Before, a read whose first or last k-mer carried a mismatch got a shortened fragment length, which biased the FLD, library-format and dovetail detection, the SAM POS and the bias coordinates. This also fixes an off-by-one in the 5' position of reverse single-end reads and orphans.ref_pos, the partner's position and the positional-bias ends now describe the fragment, as for a concordant pair.recover_mateclamps its window to the transcript (no negative window start).max_iter. It now counts that case as converged, as salmon does, and filters onalpha_out. Applied to the plain, SQUAREM and DAAREM loops.--perNucleotidePrior: usesvbPriorper nucleotide (matching the VBEM point estimate) instead ofmax(1, vbPrior), and only with VBEM.ambig_info.tsv): accumulated in u64.0Mops when a read ends exactly at the transcript start or starts exactly at its end.fld_low == 0no longer indexes position -1; the FFT path uses the same fragment lengths as the scalar loop.-gwith Ensembl GFF3: transcripts withoutgene_idare resolved to genes viaParent.--fldMean/--fldSD(finite, positive),--numGCBins/--conditionalGCBins(1..=101),--thinningFactor(>= 1) and the scoring flags are validated at parse time; salmon's negative--mpform is accepted.Performance
add_val: the sharedsum/tot_massatomics are updated once per fragment instead of once per kernel entry.Validation
cargo check --workspace --all-targets).cargo test --workspace --release,cargo fmt --checkandcargo clippy --workspace --all-targets -- -D warningspass on the last commit.🤖 Generated with Claude Code