Quickstart · Documentation · Commands · Output reference · Benchmarks
arda annotates T and B cell receptor sequences. Give it sequencing reads — bulk RNA-seq, targeted amplicon or single cell — and it returns AIRR Rearrangement records with V/D/J/C gene calls, FR1–FR4 and CDR1–CDR3 coordinates and the junction, plus a clonotype table. One command, no workflow engine, no container.
It works on nucleotide and amino-acid input, on FASTA and FASTQ, across all loci in one pass, for human and mouse (full IG + TR) and for rat, rabbit and rhesus macaque (IG only — IMGT ships no TR reference for those three). It also annotates records with no read behind them — a CDR3 amino acid plus a V and J call, as in a VDJdb row — marking up which residues each germline templates, repairing the junction, and inferring the D gene from junction length.
pip install arda-mapperThat is all. The distribution is arda-mapper and imports as arda; binary wheels carry its C++
extensions for CPython 3.10–3.13 on Linux, macOS (arm64) and Windows. The MMseqs2 search binary and
the curated germline reference are fetched automatically on first use into ~/.cache/arda — no
conda, no manual reference build. Run arda info to see what got resolved.
For a development checkout, bash setup.sh builds a uv venv, compiles the extensions and fetches
IgBLAST and MMseqs2. See the installation guide.
Pick the mode that matches your library. The mode name carries that library's speed configuration, so there are no flags to get right:
arda rnaseq --r1 R1.fq.gz --r2 R2.fq.gz -p SAMPLE -d out/ # bulk / whole-transcriptome
arda amplicon --r1 R1.fq.gz --r2 R2.fq.gz -p SAMPLE -d out/ # targeted RepSeq / 5'RACE
arda cells asm/PBMC.consensus.fq.gz -p out/PBMC # single cell, from UMI consensusEach mode runs mapping → assembly → error correction and writes:
| file | contents |
|---|---|
SAMPLE.clones.tsv |
one row per clonotype — the table most analyses start from |
SAMPLE.airr.tsv |
one AIRR Rearrangement row per mapped read |
SAMPLE.arda.json |
run report: reads in, reads mapped, why the rest were not, wall time, peak RSS |
SAMPLE.stats.tsv |
run QC in long format, for per-sample and cohort checks |
Progress goes to stderr; the output paths, one per line, to stdout.
Already have assembled sequences? Skip the pipeline:
arda annotate -i sequences.fasta -o out.airr.tsv --organism humanFull walkthrough with real output: Quickstart.
IgBLAST-quality regions, without IgBLAST at runtime. arda does the expensive work once, offline, building a reference of every in-frame V·J germline scaffold already annotated by IgBLAST. At runtime it maps a query to that reference with MMseqs2 and projects the markup through the alignment in C++. Region concordance with IgBLAST is 98–99.7 % on ~7,300 real GenBank mRNA records across five organisms. IgBLAST is needed only to rebuild the reference.
Precise on bulk RNA-seq. Over 16 datasets where every compared tool ran (5,273 real fragments, Wilson 95 % CIs): recall 0.986 [.982–.989], precision lower bound 0.889 [.881–.897]. Recall ties TRUST4; precision does not — arda's lower bound sits above every competitor's upper bound.
Cheap, not only fast. Mapping holds a flat 300–650 MB at any read depth, which is what lets many samples share a node. On a 100,000-read TRA amplicon arda uses 3.2× less peak memory and 1.4× less CPU than MiXCR, and returns the most clonotypes over the most reads.
It declines rather than guessing. A D call carries the E-value it was accepted on. A junction
whose Cys104 anchor is not in the read is not emitted. A germline with no derivable anchor is
flagged, not invented. A j_call requires J evidence rather than inheriting it from the scaffold.
Embeddable and scriptable. import arda; arda.annotate_sequences(...) returns AIRR record
dicts — no subprocess, no temp files. The CLI is plain files in, plain files out, so it drops into
any workflow engine; the same sample sheet drives the CLI, a SLURM array, the Nextflow module and
the Snakemake workflow, and a sharded run is byte-identical to a single-node one.
Every figure above is measured. Methods, repetition spread, and what each comparison is and is not entitled to claim: Benchmarks.
| Quickstart | Install, run a library, read the table that comes out |
| Choosing a mode | Bulk, amplicon or single cell, and what each writes |
| Output reference | Every file a run produces, and every column in it |
| Command reference | All 22 commands, grouped by task |
| Recipes | Copy-paste shell and polars snippets for post-run analysis |
| Use cases | Monoclonal QC, low-frequency variants, SHM, tool comparisons |
| Single cell | Reference-free per-cell contigs, chain pairing, doublets |
| Quality control | Per-sample metrics, cohort roll-up, the HTML dashboard |
| At scale | SLURM sharding, sample sheets, Nextflow and Snakemake |
| How it works | The offline reference, the runtime path, the C++ extensions |
| Background | Error correction, SHM, D segments, productivity, genotypes, recombination models |
| Glossary | Junction vs CDR3, scaffold, read group, regime, and the rest |
| API | The Python library |
import arda
records = arda.annotate_sequences(
["GACGTGCAG...", ("clone7", "CAGGTG...")], # strings or (id, seq) pairs
seqtype="nt", organism="human",
)Each record is a dict of AIRR fields — v_call, d_call/d2_call, j_call, c_call/c_class,
fwr1–fwr4, cdr1–cdr3 with 1-based closed *_start/*_end and *_aa, junction(_aa),
np1–np3, d_support, the per-segment CIGARs, v_mutations/j_mutations, productive, and
more. The TSV form passes airr.schema validation.
Reading arda's output — note that quoting is off, both when writing and when reading:
import polars as pl
clones = pl.read_csv("out/SAMPLE.clones.tsv", separator="\t", quote_char=None)
top = (clones.filter(pl.col("locus") == "TRB")
.with_columns((pl.col("duplicate_count") / pl.col("duplicate_count").sum()).alias("frequency"))
.sort("duplicate_count", descending=True))examples/ is a runnable tour, every artifact derived from real data committed to this
repository and regenerated by python examples/regenerate.py: one real mRNA per locus; the two
human reads (of 7,341, across five organisms) carrying a tandem D-D; seven VDJdb records covering
every junction-repair outcome; and a 1,035-read FASTQ that runs the whole bulk RNA-seq pipeline in
about six seconds.
bash setup.sh # uv venv, C++ extensions, IgBLAST, mmseqs
source .venv/bin/activate
python -m pytest tests/unit tests/synthetic tests/realworld -q # the CI gate
ruff check src/
make -C docs html # zero warnings requiredtests/realworld compares against IgBLAST on committed fixtures and runs offline.
tests/benchmark is opt-in with RUN_BENCHMARK=1. Install -e '.[test]' before reading a green
suite as full coverage — optional extras gate optional suites.
Layout: the package is src/arda/ (cli.py, annotate/, rnaseq/, refbuild/), the four
nanobind C++ extensions are src/_markup/, src/_prefilter/, src/_segmap/, src/_denoise/, the
committed references are in database/, and integrations/ holds the Nextflow module and Snakemake
workflow. Contributor guidance is in CLAUDE.md; design records are in
project/; planned work is in ROADMAP.md; data provenance is in
SOURCES.md; per-release changes are in CHANGELOG.md.
arda is part of the antigenomics stack and stays deliberately
inside annotation. Repertoire biology — diversity, clonality, rarefaction, overlap, cross-sample
clonotype matching — belongs to vdjtools, which takes
arda-mapper as a dependency.
GPLv3. See LICENSE.