Skip to content

Correctness and performance pass: mapping positions, EM convergence, Gibbs prior, SAM, GFF3, CLI - #1231

Open
BenjaminDEMAILLE wants to merge 13 commits into
COMBINE-lab:masterfrom
BenjaminDEMAILLE:claude/perf-and-bugfixes
Open

BenjaminDEMAILLE wants to merge 13 commits into
COMBINE-lab:masterfrom
BenjaminDEMAILLE:claude/perf-and-bugfixes

Conversation

@BenjaminDEMAILLE

@BenjaminDEMAILLE BenjaminDEMAILLE commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Rebased onto current master. Changes that master already covers (decoy block located by name, deterministic eq-class order, replicate scaling, parallel digamma and balanced EM shards, allocation-free alignment cache, posterior reuse) were dropped. One commit per change.

Correctness

  • Mapping positions: read and fragment positions now come from the projected read ends (like pufferfish's approxReadStartPos), not from the first/last seed. Before, a read whose first or last k-mer carried a mismatch got a shortened fragment length, which biased the FLD, library-format and dovetail detection, the SAM POS and the bias coordinates. This also fixes an off-by-one in the 5' position of reverse single-end reads and orphans.
  • Recovered orphan pairs: ref_pos, the partner's position and the positional-bias ends now describe the fragment, as for a concordant pair. recover_mate clamps its window to the transcript (no negative window start).
  • Paired seq-bias (ISR and similar): the 5'/3' contexts were attributed to the wrong end when mate 1 is the reverse mate.
  • Refseq store: reference lengths are checked against the index (duplicate names no longer write a corrupt store).
  • EM/VBEM: with no transcript above the convergence cutoff (e.g. no fragments), the loop ran to max_iter. It now counts that case as converged, as salmon does, and filters on alpha_out. Applied to the plain, SQUAREM and DAAREM loops.
  • Gibbs prior under --perNucleotidePrior: uses vbPrior per nucleotide (matching the VBEM point estimate) instead of max(1, vbPrior), and only with VBEM.
  • Ambiguity counts (ambig_info.tsv): accumulated in u64.
  • SAM/BAM: CIGARs no longer contain zero-length 0M ops when a read ends exactly at the transcript start or starts exactly at its end.
  • Bias convolutions: fld_low == 0 no longer indexes position -1; the FFT path uses the same fragment lengths as the scalar loop.
  • -g with Ensembl GFF3: transcripts without gene_id are resolved to genes via Parent.
  • CLI: --fldMean/--fldSD (finite, positive), --numGCBins/--conditionalGCBins (1..=101), --thinningFactor (>= 1) and the scoring flags are validated at parse time; salmon's negative --mp form is accepted.

Performance

  • Gibbs: multinomial buffers are reused across rounds instead of allocated per class.
  • FLD add_val: the shared sum/tot_mass atomics are updated once per fragment instead of once per kernel entry.

Validation

  • Every commit builds (cargo check --workspace --all-targets).
  • cargo test --workspace --release, cargo fmt --check and cargo clippy --workspace --all-targets -- -D warnings pass on the last commit.
  • New regression tests: projected fragment length, empty EM convergence, parallel vs sequential EM, CIGAR overhangs, Ensembl GFF3.

🤖 Generated with Claude Code

@BenjaminDEMAILLE
BenjaminDEMAILLE force-pushed the claude/perf-and-bugfixes branch from 4ae22ba to 907e832 Compare October 1, 2026 06:26
@BenjaminDEMAILLE
BenjaminDEMAILLE marked this pull request as ready for review October 1, 2026 06:27
BenjaminDEMAILLE and others added 13 commits October 1, 2026 08:31
Positions were taken from the first/last seed of the MEM chain, so an
unmatched read prefix/suffix (e.g. a mismatch in the edge k-mer) shortened
fragments and shifted coordinates. MappingCandidate now records the read
length and exposes proj_start()/proj_end() (chain start minus unmatched
prefix, chain end plus unmatched suffix), used for fragment lengths (FLD),
library-format and dovetail detection, SAM POS and bias 5' coordinates.
Also fixes the off-by-one 5' position for reverse single-end reads and
orphans (last base is proj_end() - 1).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A pair recovered from an orphan anchor now gets the fragment's leftmost
base as ref_pos, the partner's real leftmost position (a reverse anchor
ends the fragment, so the partner starts it) and fragment 5'/3' ends for
positional bias, as for a concordant pair. recover_mate clamps its search
window to the transcript (no negative window start when the read
overhangs) and reports the fragment length relative to the projected
anchor start.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
For a paired fragment ref_pos is the fragment's leftmost base whichever
mate is forward, so the forward 5' context is ref_pos and the reverse 5'
context is the fragment's last base. The old code branched on mate 1's
orientation and mis-attributed contexts for ISR-type libraries where
mate 1 is the reverse read.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Fail when a reference's FASTA length differs from its length in the index
(e.g. duplicate reference names) instead of writing a corrupt store.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max_rel_diff now filters on alpha_out (as salmon does) and returns -inf
when no transcript clears the cutoff, which counts as converged. Before,
an input with no fragments ran EM/VBEM for max_iter iterations.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rNucleotidePrior

The Gibbs prior kept max(1, vbPrior) when switched to per-nucleotide mode,
so it no longer matched the VBEM point estimate (vbPrior * max(1, effLen)),
and plain EM runs switched to a per-nucleotide 1e-3 prior. The
per-nucleotide prior now applies only with VBEM and uses vbPrior itself.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ambig_info.tsv counts were summed in u32 after truncating the u64 class
counts, which can overflow on deep libraries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A read ending exactly at the transcript start or starting exactly at its
end produced a 0M op (rejected by htslib/samtools); it is now written as
all soft-clip. Adds unit tests for overhang_cigar.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A zero-length fragment has no end base; starting the convolution at
length 0 indexed position -1 and panicked.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Ensembl GFF3 puts gene_id only on gene features and links transcripts to
them with Parent=<gene ID>. Transcripts without gene_id are now resolved
through their Parent after the pass, falling back to the bare identifier
of gene:ENSG... IDs when the gene feature is missing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Reject non-positive or non-finite --fldMean/--fldSD, GC bin counts
outside 1..=101, --thinningFactor 0 and out-of-range scoring flags at
parse time. --mp accepts salmon's negative convention (--mp -4); its
magnitude is used.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each Gibbs chain keeps one probs/draws scratch pair for all its rounds
instead of allocating a draws vector per multi-transcript class per round
(and a probs vector per round). Also adds a test that parallel and
sequential EM/VBEM agree.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
add_val accumulates the fragment's total and length-weighted mass locally
and publishes them with one atomic update each, instead of a CAS loop on
the shared sum/tot_mass per kernel entry, avoiding cache-line contention
between workers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@BenjaminDEMAILLE
BenjaminDEMAILLE force-pushed the claude/perf-and-bugfixes branch from 907e832 to 5053d6f Compare October 1, 2026 09:15
@BenjaminDEMAILLE BenjaminDEMAILLE changed the title Correctness and performance pass: mapping positions, decoy index, EM, uncertainty, CLI Oct 1, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

1 participant