A gate for biological inputs. Validate gene symbols, ontology terms, variant formats, and database identifiers, the same way, with the same answer, in R, Python, and JavaScript.
Documentation: R package (pkgdown), Python package (MkDocs), and JavaScript package (TypeDoc), from one landing page.
biobouncer lives in a monorepo: the R package in pkg-r/, the Python package in
pkg-py/, and the JavaScript package in pkg-js/. This matters only when
installing from GitHub, where you point the installer at the subdirectory.
Python
pip install biobouncer
# development version from GitHub (package is in the pkg-py/ subdirectory)
pip install "git+https://github.com/samuelbharti/biobouncer.git#subdirectory=pkg-py"R
Install from CRAN:
install.packages("biobouncer")
# development version from GitHub (package is in the pkg-r/ subdirectory)
pak::pak("samuelbharti/biobouncer/pkg-r")R-universe also serves prebuilt binaries of the latest release:
install.packages("biobouncer", repos = "https://samuelbharti.r-universe.dev").
JavaScript
npm install biobouncerThe package ships a Node build and a browser build. See the JavaScript docs for the runtime targets and the async entry points.
If you build analyses or Shiny/Dash apps in computational biology, you keep rewriting the same guards: is this a real gene symbol? a well-formed MONDO id? a valid HGVS string? a UniProt accession that actually exists? Those checks end up scattered across projects as ad-hoc regexes and utility functions, and the R version and the Python version quietly disagree on edge cases.
biobouncer puts those checks in one place, behind one small API, and guarantees
that R, Python, and JavaScript give the same verdict for the same input by
testing all three against a shared conformance corpus. It does not try to replace annotation
engines like biomaRt, ensembldb, or mygene. It validates inputs before they
reach those tools.
- One entry point, many sources.
check_id()/is_valid_id()work across 50 databases and ontologies (MONDO, EFO, HGNC, Ensembl, RefSeq, dbSNP, UniProt, ChEBI, GO, HGVS, and more) selected with a singlesource_dbargument. - Checking modes. Choose how strict and how online you want to be:
pattern(offline regex/grammar),cache(offline existence against a pinned snapshot),remote(live existence against the source's API), orexistence(the snapshot if one answers, else remote, else pattern). - Species-, source-, and version-aware. Ask not just "is this valid?" but "was this valid for this species, in this source, at this version?"
- Rich, vectorized results. Get a per-element table of
valid/normalized/suggestion, not just a single boolean, so you can filter or repair a column, not just reject one field. - Plugs into the tools you already use. Adapters for
pandera,pydantic,great_expectations, andnarwhalsin Python, andshinyvalidate,checkmate, andassertr/validate/pointblankin R. - Works from the shell. The Python package installs a
biobouncercommand that validates ids from a file or a pipe and exits non-zero on any invalid input, for use in scripts and CI. - Reproducible by design.
patternandcachemodes are pure functions of pinned data; every result records the snapshot version it came from.
demo/ has two notebooks, one in
Python and one in R,
and a JavaScript script. All three run the same story
over the same messy data so you can see the packages reach the same answers.
The notebooks cover all four modes and end with the framework adapters; the
script covers pattern, cache, and remote. All run offline.
biobouncer publishes its docs in the llms.txt format, so a coding agent can read the whole API in one pass:
- llms.txt: a short index of links.
- llms-full.txt: the full API and usage in one file.
Point your agent (Claude Code, Cursor, and similar) at llms-full.txt, then ask
it to validate or clean a data file. It can install biobouncer with the same pip
or install.packages() commands above. For example:
Read https://www.samuelbharti.com/biobouncer/llms-full.txt. Then use biobouncer in Python to validate and repair the
genecolumn ofdata.csvagainsthgncin cache mode, write the cleaned table todata_clean.csv, and preserve the row order.
R
library(biobouncer)
# 1. pattern mode: offline, deterministic, no reference data
is_valid_id("MONDO:0005148", source_db = "mondo", how = "pattern")
#> [1] TRUE
# 2. cache mode: existence against a pinned local snapshot
is_valid_id("MONDO:0005148", source_db = "mondo", how = "cache",
version = "sample")
#> [1] TRUE
# 3. remote mode: live existence check against the source
is_valid_id("ENSG00000139618", source_db = "ensembl", how = "remote",
species = "homo_sapiens")
#> [1] TRUE
# Rich, vectorized result over a whole column
check_id(
c("MONDO:0005148", "MONDO:9999999", "mondo:5148"),
source_db = "mondo",
how = "cache",
version = "sample"
)
#> # A tibble: 3 x 9
#> input valid normalized suggestion source_db version species how error
#> <chr> <lgl> <chr> <chr> <chr> <chr> <chr> <chr> <chr>
#> 1 MONDO:0005148 TRUE MONDO:0005148 NA mondo sample NA cache NA
#> 2 MONDO:9999999 FALSE NA NA mondo sample NA cache NA
#> 3 mondo:5148 FALSE NA MONDO:0005148 mondo sample NA cache NAPython
import biobouncer as bg
bg.is_valid_id("MONDO:0005148", source_db="mondo", how="pattern")
# True
bg.is_valid_id(
"ENSG00000139618", source_db="ensembl", how="remote",
species="homo_sapiens",
)
# True
# Vectorized: returns a list of per-element Result records, in input order
bg.check_id(
["MONDO:0005148", "MONDO:9999999", "mondo:5148"],
source_db="mondo", how="cache", version="sample",
)JavaScript
import { checkId, isValidId, isValidIdAsync } from "biobouncer";
// 1. pattern mode: offline, deterministic, and synchronous
isValidId("MONDO:0005148", "mondo"); // true
// 2. cache mode: existence against a pinned local snapshot
isValidId("MONDO:0005148", "mondo", { how: "cache", version: "sample" }); // true
// 3. remote mode reaches the network, so it takes the async entry point
await isValidIdAsync("ENSG00000139618", "ensembl", {
how: "remote",
species: "homo_sapiens",
}); // true
// Rich, per-item results over a whole column, in input order
checkId(["MONDO:0005148", "MONDO:9999999", "mondo:5148"], "mondo", {
how: "cache",
version: "sample",
});The everyday job is a whole column: which values are wrong, and can you fix the
ones you can. report_id() / report() validate the column and print a summary;
repair_id() / Report.repair() substitute the fixable values (a withdrawn gene
symbol becomes its successor) and leave valid, unmappable, and missing values
untouched. In cache mode the snapshot version defaults to the latest installed,
so you do not have to name one.
R
genes <- c("TP53", "MLL", "notagene", NA)
report_id(genes, "hgnc", how = "cache")
#> # biobouncer report on hgnc (cache mode): 1 valid, 1 repairable, 1 invalid, 1 missing of 4
#> # A tibble: 4 x 9
#> input valid normalized suggestion source_db version species how error
#> <chr> <lgl> <chr> <chr> <chr> <chr> <chr> <chr> <chr>
#> 1 TP53 TRUE TP53 NA hgnc 2026-07-07 NA cache NA
#> 2 MLL FALSE NA KMT2A hgnc 2026-07-07 NA cache NA
#> 3 notagene FALSE NA NA hgnc 2026-07-07 NA cache NA
#> 4 NA NA NA NA hgnc 2026-07-07 NA cache NA
repair_id(genes, "hgnc", how = "cache")
#> [1] "TP53" "KMT2A" "notagene" NAPython
genes = ["TP53", "MLL", "notagene", None]
rep = bg.report(genes, "hgnc", how="cache")
rep
# <biobouncer report on 'hgnc' (cache mode): 1 valid, 1 repairable, 1 invalid, 1 missing of 4>
rep.repair()
# ['TP53', 'KMT2A', 'notagene', None]JavaScript
import { report } from "biobouncer";
const genes = ["TP53", "MLL", "notagene", null];
const rep = report(genes, "hgnc", { how: "cache" });
rep.summary;
// { total: 4, valid: 1, invalid: 1, repairable: 1, missing: 1, indeterminate: 0 }
rep.repair();
// ["TP53", "KMT2A", "notagene", null]report/report_id are for inspecting and cleaning; to enforce validity inside
a framework (pandera, Great Expectations, pydantic, shiny) use the adapters.
| Mode | What it answers | Network | Reproducible | Speed |
|---|---|---|---|---|
pattern |
Is the string well-formed for this source? | no | yes | fast |
cache |
Does the id exist in a pinned local snapshot? | no | yes | fast |
remote |
Does the id exist right now in the source? | yes | no | slow |
how = "existence" is a convenience that tries cache first and falls back to
remote if no local snapshot is available.
The split matters for reproducibility: pattern and cache are pure functions
of code and pinned data, so the same call always returns the same answer.
remote reflects the live source and can change between runs; every result
records which mode and snapshot produced it.
Identifiers are not valid in a vacuum. A symbol can be current in one species
and meaningless in another; an id can exist in one release of a source and be
retired in the next. biobouncer makes these explicit arguments:
# Same accession, different species contexts
check_id("ENSMUSG00000059552", source_db = "ensembl",
species = "mus_musculus", how = "remote") # valid
check_id("ENSMUSG00000059552", source_db = "ensembl",
species = "homo_sapiens", how = "remote") # not a human id
# Retired symbols map to their approved successor
check_id("MLL", source_db = "hgnc", how = "cache", version = "sample")
#> valid = FALSE, suggestion = "KMT2A" (MLL was renamed KMT2A)speciesaccepts a name ("homo_sapiens") or an NCBI Taxonomy id (9606). It is enforced by the sources for which it applies (such as Ensembl and UniProt) and ignored by the rest.versionselects the snapshot forcachemode. The packages ship a smallsamplesnapshot;biobouncer_pull()fetches full, dated snapshots for the OBO ontologies.
Every check_id() row carries enough context to be self-describing:
| column | meaning |
|---|---|
input |
the original value, unchanged |
valid |
logical verdict |
normalized |
canonical form when valid (e.g. case/prefix normalized) |
suggestion |
best-effort correction when invalid but mappable |
source_db |
source the check ran against |
version |
snapshot/release that produced the answer |
species |
species context, when applicable |
how |
mode used (pattern / cache / remote / existence) |
error |
reason a remote check was left indeterminate, else NA/None |
biobouncer checks 50 sources. A selection is shown below. Run source_info() for
the full list, or read the sources cookbook for
R or
Python, which gives an
example id and the modes each source supports.
source_db |
Source | Example id | pattern | cache | remote | species-aware |
|---|---|---|---|---|---|---|
mondo |
MONDO disease ontology | MONDO:0005148 |
✓ | ✓ | ✓ | - |
efo |
Experimental Factor Ont. | EFO:0000400 |
✓ | ✓ | ✓ | - |
go |
Gene Ontology terms | GO:0006915 |
✓ | ✓ | ✓ | - |
chebi |
ChEBI compounds | CHEBI:15377 |
✓ | ✓ | ✓ | - |
hgnc |
HGNC gene symbols | TP53 |
~ | ✓ | ✓ | - |
ensembl |
Ensembl gene/transcript | ENSG00000139618 |
✓ | - | ✓ | ✓ |
opentargets |
Open Targets targets | ENSG00000139618 |
✓ | - | ✓ | - |
refseq |
RefSeq accessions | NM_000546.6 |
✓ | - | ✓ | - |
uniprot |
UniProt accessions | P04637 |
✓ | - | ✓ | ✓ |
dbsnp |
dbSNP variants | rs7412 |
✓ | - | ✓ | - |
hgvs |
HGVS variant syntax | NM_004006.2:c.4375C>T |
✓† | - | ✓ | - |
gnomad |
gnomAD variant coordinate | 1-55516888-G-A |
✓ | - | - | - |
mgi |
MGI mouse accessions | MGI:97306 |
✓ | - | - | - |
rgd |
RGD rat identifiers | RGD:3059 |
✓ | - | - | - |
✓ supported · ~ shape check only, a loose token match · - not available ·
† syntax only, a single regex (not coordinate-level validation). The Open
Targets connector checks whether a human Ensembl gene id is a target the platform
covers, through its GraphQL API. Identifier patterns come from the
Identifiers.org / Bioregistry registries where available.
biobouncer provides the domain checks; your existing validation framework
provides the plumbing.
pandera (Python)
import pandera.pandas as pa
from biobouncer.checks import is_id
schema = pa.DataFrameSchema({
"disease_id": pa.Column(str, is_id(source_db="mondo", how="cache",
version="sample")),
"target_id": pa.Column(str, is_id(source_db="ensembl",
species="homo_sapiens")),
})pydantic (Python)
from pydantic import BaseModel
from biobouncer.types import Id
class Association(BaseModel):
disease: Id("mondo", how="pattern")
target: Id("ensembl", species="homo_sapiens", how="pattern")shinyvalidate (R)
iv <- InputValidator$new()
iv$add_rule("gene", sv_biobouncer(source_db = "hgnc", how = "cache", version = "sample"))
iv$enable()checkmate / assertr (R)
# stop early in a pipeline
assert_valid_id(df$disease, source_db = "mondo", how = "cache", version = "sample")
# assertr verb inside a dplyr chain
df |> assertr::verify(
is_valid_id(disease, source_db = "mondo", how = "cache", version = "sample")
)Offline cache mode reads versioned snapshots of source identifier sets.
Snapshots are pinned (never auto-updated silently) so analyses stay
reproducible, and are distributed separately from the code so they can be
refreshed without a package release:
biobouncer_snapshots() # list installed snapshots
biobouncer_pull("mondo") # fetch the current MONDO snapshot
biobouncer_cache_dir() # where snapshots liveremote mode caches responses locally and respects each source's rate limits.
Adding a source should be small and declarative: a prefix, a pattern, optional
species/version metadata, and (optionally) a cache builder and a remote
resolver. See CONTRIBUTING.md and the source-registry spec in PLAN.md.
Barret Schloerke and Carson Sievert advise this work as thesis advisors. Posit Software, PBC funded early work on this package and holds copyright together with the author.
MIT © biobouncer authors.
If biobouncer supports your work, please cite it. A preprint is in preparation,
and its citation will be added here as soon as it is available. Until then, cite
the software archived on Zenodo:
Bharti, S. (2026). biobouncer: Validate Biological Identifiers and Inputs. Zenodo. https://doi.org/10.5281/zenodo.21346522
The DOI above always resolves to the latest release; each release also has its
own version DOI on Zenodo. Machine-readable metadata is in CITATION.cff.
