Skip to content

Tags: antigenomics/arda

Tags

v2.36.0

Toggle v2.36.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
2.36.0 — a junction naming neither side gets a locus proposed too (#146)

2.34.0 proposed the missing side from the junction, but only within the locus
the *other* side named. A record naming neither came back refused: no locus, no
boundary, nothing for a nucleotide stage to work on.

No new rule. Every locus the organism ships now competes, scored by the same
`anchor_depth` from each germline's own anchor, and a locus only wins by
explaining residues at BOTH ends -- which is what keeps a TRA junction out of
TRD, where the V genes are shared and the J genes are not.

522 of VDJdb's curated `chunks` records name neither side (461 distinct keys).
All 461 get a locus and 459 come back `good`, against none before. The proposed
locus agrees with the `cdr3.alpha`/`cdr3.beta` column the record was filed under
on 457 of 461 (99.13 %) -- 318 TRB, 139 TRA -- and all four disagreements are
`CACD...DKLIF`: TRDV2's own anchor and TRDJ1's own ending, in a schema with no
delta column.

Each end must explain its own anchor residue or no locus is named. Without that
floor `QQQQQQQQQQQQ` agreed with nothing on every locus and was still handed
TRA; over the 461 real keys the winning locus clears it on every one.

scripts/audit_cdr3fix.py kept blank-side keys out. It required both a V and a J
to be named, so the A/B instrument was blind to exactly the rows 2.34.0 and this
release change. Corpus 189,596 -> 192,726 keys, digest rebases to
0df8541ed1060fe1; on the 189,596 the old filter kept, the digest is unchanged at
2b75491b380310aa, so nothing that already worked moved.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

v2.35.0

Toggle v2.35.0's commit message
Point the D-posterior references at what actually answers that questi…

…on now

vdjtools 4.8.0 ships one D estimator, not two: the model names the gene from its own scenario
weights and arda places it, so `vdjtools.model.posterior_d` / `posterior_d_batch` / `load_d_prior`
do not exist. Eight places here still named them -- `scenarios.py`, `cli.py` (two help strings and
a docstring), `docs/api.rst`, `docs/scenarios.rst`, `docs/reference_build.rst`, and two skill
reference pages -- and a reader following any of them lands on an ImportError.

They now point at `vdjtools.model.annotate_junctions` where the question is the amino-acid one, and
at `arda.hmm.posterior_d` / `arda.hmm.model_for(prior=)` where it is the nucleotide one, which is
also the correction to `PriorTable`'s docstring: both consumers of `d_prior.tsv` are in arda, and
one of them is the SHM-aware scorer B cells need.

Docs build clean with -W.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

v2.34.1

Toggle v2.34.1's commit message
2.34.1 — walk a germline run once

Packaging release for 67de699 and 0c51226: `_extend` goes from three passes over a germline run to
one (16,287 -> 18,338 keys/s over VDJdb's 189,596-key corpus, audit digest `2b75491b380310aa`
unchanged), and the two invariants that used to hold only by construction are tests -- every residue
a boundary credits is germline of the allele it names, and a functional segment carries its canonical
anchor with `TRAJ35*01` named as the single real exception.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

v2.34.0

Toggle v2.34.0's commit message
2.34.0 — propose a blank V or J from the junction instead of refusing…

… the record

A submission is allowed to leave one side out, and a blank is not a reason to refuse: the locus
comes from the side that IS named, and the junction is evidence about the missing one. New
`Cdr3Markup.proposed` says which side was never curated — a different fact from `allele`, which
means the submission named a different allele of the same gene.

Over VDJdb's 192,726 distinct curation keys, 3,130 leave a side blank (644 no V, 2,947 no J).
2,532 now get a segment, 2,504 of them `good` with both boundaries placed; the 598 that stay
refused name neither side, so there is no locus to propose within. Before this every one of the
3,130 came back `FailedBadSegment` with no boundary and no repair, which is why the consumer
carried its own segment proposer (`vdjdb.annotate.segments`, 185 lines) — that module can go.

ABSENT and WRONG stay different. `TRBVnope*01` keeps its `FailedBadSegment`: a submission that
names something wrong has a defect a curator must see, where one that names nothing has a gap
the junction can fill. Proposing for both would hide the first inside the second.

The proposal reuses the re-call machinery with `min_gain=0` — there is no curator's call to beat
and none to protect, so the whole locus competes from the junction's own start and the margin
becomes the floor on the evidence needed to name a segment at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

v2.33.0

Toggle v2.33.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
Merge pull request #145 from antigenomics/feature/kmer-cdr3fix

2.33.0 — the repair policy, and the D posterior moves to vdjtools

v2.31.0

Toggle v2.31.0's commit message
arda-mapper 2.31.0: the germline boundary in nucleotides

v2.30.1

Toggle v2.30.1's commit message
arda 2.30.1

v2.30.0

Toggle v2.30.0's commit message
arda 2.30.0

v2.29.0

Toggle v2.29.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
Merge pull request #130 from antigenomics/release/2.29.0

release: 2.29.0

v2.28.0

Toggle v2.28.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
Merge pull request #127 from antigenomics/release/2.28.0

release: 2.28.0