Skip to content

Repository files navigation

biobouncer biobouncer logo

Lifecycle: stable CRAN status CRAN downloads r-universe npm PyPI DOI License: MIT

A gate for biological inputs. Validate gene symbols, ontology terms, variant formats, and database identifiers, the same way, with the same answer, in R, Python, and JavaScript.

biobouncer takes a messy column of biological identifiers, checks each through one gate with four modes (pattern, cache, remote, existence), and returns one labeled verdict (valid, repairable, invalid, or missing) that is the same in R, Python, and JavaScript.

Documentation: R package (pkgdown), Python package (MkDocs), and JavaScript package (TypeDoc), from one landing page.


Installation

biobouncer lives in a monorepo: the R package in pkg-r/, the Python package in pkg-py/, and the JavaScript package in pkg-js/. This matters only when installing from GitHub, where you point the installer at the subdirectory.

Python

pip install biobouncer

# development version from GitHub (package is in the pkg-py/ subdirectory)
pip install "git+https://github.com/samuelbharti/biobouncer.git#subdirectory=pkg-py"

R

Install from CRAN:

install.packages("biobouncer")

# development version from GitHub (package is in the pkg-r/ subdirectory)
pak::pak("samuelbharti/biobouncer/pkg-r")

R-universe also serves prebuilt binaries of the latest release: install.packages("biobouncer", repos = "https://samuelbharti.r-universe.dev").

JavaScript

npm install biobouncer

The package ships a Node build and a browser build. See the JavaScript docs for the runtime targets and the async entry points.

Motivation

If you build analyses or Shiny/Dash apps in computational biology, you keep rewriting the same guards: is this a real gene symbol? a well-formed MONDO id? a valid HGVS string? a UniProt accession that actually exists? Those checks end up scattered across projects as ad-hoc regexes and utility functions, and the R version and the Python version quietly disagree on edge cases.

biobouncer puts those checks in one place, behind one small API, and guarantees that R, Python, and JavaScript give the same verdict for the same input by testing all three against a shared conformance corpus. It does not try to replace annotation engines like biomaRt, ensembldb, or mygene. It validates inputs before they reach those tools.

Features

  • One entry point, many sources. check_id() / is_valid_id() work across 50 databases and ontologies (MONDO, EFO, HGNC, Ensembl, RefSeq, dbSNP, UniProt, ChEBI, GO, HGVS, and more) selected with a single source_db argument.
  • Checking modes. Choose how strict and how online you want to be: pattern (offline regex/grammar), cache (offline existence against a pinned snapshot), remote (live existence against the source's API), or existence (the snapshot if one answers, else remote, else pattern).
  • Species-, source-, and version-aware. Ask not just "is this valid?" but "was this valid for this species, in this source, at this version?"
  • Rich, vectorized results. Get a per-element table of valid / normalized / suggestion, not just a single boolean, so you can filter or repair a column, not just reject one field.
  • Plugs into the tools you already use. Adapters for pandera, pydantic, great_expectations, and narwhals in Python, and shinyvalidate, checkmate, and assertr/validate/pointblank in R.
  • Works from the shell. The Python package installs a biobouncer command that validates ids from a file or a pipe and exits non-zero on any invalid input, for use in scripts and CI.
  • Reproducible by design. pattern and cache modes are pure functions of pinned data; every result records the snapshot version it came from.

Demos

demo/ has two notebooks, one in Python and one in R, and a JavaScript script. All three run the same story over the same messy data so you can see the packages reach the same answers. The notebooks cover all four modes and end with the framework adapters; the script covers pattern, cache, and remote. All run offline.

AI agents

biobouncer publishes its docs in the llms.txt format, so a coding agent can read the whole API in one pass:

Point your agent (Claude Code, Cursor, and similar) at llms-full.txt, then ask it to validate or clean a data file. It can install biobouncer with the same pip or install.packages() commands above. For example:

Read https://www.samuelbharti.com/biobouncer/llms-full.txt. Then use biobouncer in Python to validate and repair the gene column of data.csv against hgnc in cache mode, write the cleaned table to data_clean.csv, and preserve the row order.

Usage

R

library(biobouncer)

# 1. pattern mode: offline, deterministic, no reference data
is_valid_id("MONDO:0005148", source_db = "mondo", how = "pattern")
#> [1] TRUE

# 2. cache mode: existence against a pinned local snapshot
is_valid_id("MONDO:0005148", source_db = "mondo", how = "cache",
            version = "sample")
#> [1] TRUE

# 3. remote mode: live existence check against the source
is_valid_id("ENSG00000139618", source_db = "ensembl", how = "remote",
            species = "homo_sapiens")
#> [1] TRUE

# Rich, vectorized result over a whole column
check_id(
  c("MONDO:0005148", "MONDO:9999999", "mondo:5148"),
  source_db = "mondo",
  how       = "cache",
  version   = "sample"
)
#> # A tibble: 3 x 9
#>   input         valid normalized    suggestion    source_db version species how   error
#>   <chr>         <lgl> <chr>         <chr>         <chr>     <chr>   <chr>   <chr> <chr>
#> 1 MONDO:0005148 TRUE  MONDO:0005148 NA            mondo     sample  NA      cache NA
#> 2 MONDO:9999999 FALSE NA            NA            mondo     sample  NA      cache NA
#> 3 mondo:5148    FALSE NA            MONDO:0005148 mondo     sample  NA      cache NA

Python

import biobouncer as bg

bg.is_valid_id("MONDO:0005148", source_db="mondo", how="pattern")
# True

bg.is_valid_id(
    "ENSG00000139618", source_db="ensembl", how="remote",
    species="homo_sapiens",
)
# True

# Vectorized: returns a list of per-element Result records, in input order
bg.check_id(
    ["MONDO:0005148", "MONDO:9999999", "mondo:5148"],
    source_db="mondo", how="cache", version="sample",
)

JavaScript

import { checkId, isValidId, isValidIdAsync } from "biobouncer";

// 1. pattern mode: offline, deterministic, and synchronous
isValidId("MONDO:0005148", "mondo"); // true

// 2. cache mode: existence against a pinned local snapshot
isValidId("MONDO:0005148", "mondo", { how: "cache", version: "sample" }); // true

// 3. remote mode reaches the network, so it takes the async entry point
await isValidIdAsync("ENSG00000139618", "ensembl", {
  how: "remote",
  species: "homo_sapiens",
}); // true

// Rich, per-item results over a whole column, in input order
checkId(["MONDO:0005148", "MONDO:9999999", "mondo:5148"], "mondo", {
  how: "cache",
  version: "sample",
});

Cleaning a column

The everyday job is a whole column: which values are wrong, and can you fix the ones you can. report_id() / report() validate the column and print a summary; repair_id() / Report.repair() substitute the fixable values (a withdrawn gene symbol becomes its successor) and leave valid, unmappable, and missing values untouched. In cache mode the snapshot version defaults to the latest installed, so you do not have to name one.

R

genes <- c("TP53", "MLL", "notagene", NA)

report_id(genes, "hgnc", how = "cache")
#> # biobouncer report on hgnc (cache mode): 1 valid, 1 repairable, 1 invalid, 1 missing of 4
#> # A tibble: 4 x 9
#>   input    valid normalized suggestion source_db version    species how   error
#>   <chr>    <lgl> <chr>      <chr>      <chr>     <chr>      <chr>   <chr> <chr>
#> 1 TP53     TRUE  TP53       NA         hgnc      2026-07-07 NA      cache NA
#> 2 MLL      FALSE NA         KMT2A      hgnc      2026-07-07 NA      cache NA
#> 3 notagene FALSE NA         NA         hgnc      2026-07-07 NA      cache NA
#> 4 NA       NA    NA         NA         hgnc      2026-07-07 NA      cache NA

repair_id(genes, "hgnc", how = "cache")
#> [1] "TP53"     "KMT2A"    "notagene" NA

Python

genes = ["TP53", "MLL", "notagene", None]

rep = bg.report(genes, "hgnc", how="cache")
rep
# <biobouncer report on 'hgnc' (cache mode): 1 valid, 1 repairable, 1 invalid, 1 missing of 4>

rep.repair()
# ['TP53', 'KMT2A', 'notagene', None]

JavaScript

import { report } from "biobouncer";

const genes = ["TP53", "MLL", "notagene", null];

const rep = report(genes, "hgnc", { how: "cache" });
rep.summary;
// { total: 4, valid: 1, invalid: 1, repairable: 1, missing: 1, indeterminate: 0 }

rep.repair();
// ["TP53", "KMT2A", "notagene", null]

report/report_id are for inspecting and cleaning; to enforce validity inside a framework (pandera, Great Expectations, pydantic, shiny) use the adapters.

Modes

Mode What it answers Network Reproducible Speed
pattern Is the string well-formed for this source? no yes fast
cache Does the id exist in a pinned local snapshot? no yes fast
remote Does the id exist right now in the source? yes no slow

how = "existence" is a convenience that tries cache first and falls back to remote if no local snapshot is available.

The split matters for reproducibility: pattern and cache are pure functions of code and pinned data, so the same call always returns the same answer. remote reflects the live source and can change between runs; every result records which mode and snapshot produced it.

Species and versions

Identifiers are not valid in a vacuum. A symbol can be current in one species and meaningless in another; an id can exist in one release of a source and be retired in the next. biobouncer makes these explicit arguments:

# Same accession, different species contexts
check_id("ENSMUSG00000059552", source_db = "ensembl",
         species = "mus_musculus",  how = "remote")   # valid
check_id("ENSMUSG00000059552", source_db = "ensembl",
         species = "homo_sapiens", how = "remote")    # not a human id

# Retired symbols map to their approved successor
check_id("MLL", source_db = "hgnc", how = "cache", version = "sample")
#> valid = FALSE, suggestion = "KMT2A"  (MLL was renamed KMT2A)
  • species accepts a name ("homo_sapiens") or an NCBI Taxonomy id (9606). It is enforced by the sources for which it applies (such as Ensembl and UniProt) and ignored by the rest.
  • version selects the snapshot for cache mode. The packages ship a small sample snapshot; biobouncer_pull() fetches full, dated snapshots for the OBO ontologies.

Result schema

Every check_id() row carries enough context to be self-describing:

column meaning
input the original value, unchanged
valid logical verdict
normalized canonical form when valid (e.g. case/prefix normalized)
suggestion best-effort correction when invalid but mappable
source_db source the check ran against
version snapshot/release that produced the answer
species species context, when applicable
how mode used (pattern / cache / remote / existence)
error reason a remote check was left indeterminate, else NA/None

Sources

biobouncer checks 50 sources. A selection is shown below. Run source_info() for the full list, or read the sources cookbook for R or Python, which gives an example id and the modes each source supports.

source_db Source Example id pattern cache remote species-aware
mondo MONDO disease ontology MONDO:0005148 ✓ ✓ ✓ -
efo Experimental Factor Ont. EFO:0000400 ✓ ✓ ✓ -
go Gene Ontology terms GO:0006915 ✓ ✓ ✓ -
chebi ChEBI compounds CHEBI:15377 ✓ ✓ ✓ -
hgnc HGNC gene symbols TP53 ~ ✓ ✓ -
ensembl Ensembl gene/transcript ENSG00000139618 ✓ - ✓ ✓
opentargets Open Targets targets ENSG00000139618 ✓ - ✓ -
refseq RefSeq accessions NM_000546.6 ✓ - ✓ -
uniprot UniProt accessions P04637 ✓ - ✓ ✓
dbsnp dbSNP variants rs7412 ✓ - ✓ -
hgvs HGVS variant syntax NM_004006.2:c.4375C>T ✓† - ✓ -
gnomad gnomAD variant coordinate 1-55516888-G-A ✓ - - -
mgi MGI mouse accessions MGI:97306 ✓ - - -
rgd RGD rat identifiers RGD:3059 ✓ - - -

✓ supported · ~ shape check only, a loose token match · - not available · † syntax only, a single regex (not coordinate-level validation). The Open Targets connector checks whether a human Ensembl gene id is a target the platform covers, through its GraphQL API. Identifier patterns come from the Identifiers.org / Bioregistry registries where available.

Validation frameworks

biobouncer provides the domain checks; your existing validation framework provides the plumbing.

pandera (Python)

import pandera.pandas as pa
from biobouncer.checks import is_id

schema = pa.DataFrameSchema({
    "disease_id": pa.Column(str, is_id(source_db="mondo", how="cache",
                                       version="sample")),
    "target_id":  pa.Column(str, is_id(source_db="ensembl",
                                       species="homo_sapiens")),
})

pydantic (Python)

from pydantic import BaseModel
from biobouncer.types import Id

class Association(BaseModel):
    disease: Id("mondo", how="pattern")
    target:  Id("ensembl", species="homo_sapiens", how="pattern")

shinyvalidate (R)

iv <- InputValidator$new()
iv$add_rule("gene", sv_biobouncer(source_db = "hgnc", how = "cache", version = "sample"))
iv$enable()

checkmate / assertr (R)

# stop early in a pipeline
assert_valid_id(df$disease, source_db = "mondo", how = "cache", version = "sample")

# assertr verb inside a dplyr chain
df |> assertr::verify(
  is_valid_id(disease, source_db = "mondo", how = "cache", version = "sample")
)

Caching

Offline cache mode reads versioned snapshots of source identifier sets. Snapshots are pinned (never auto-updated silently) so analyses stay reproducible, and are distributed separately from the code so they can be refreshed without a package release:

biobouncer_snapshots()          # list installed snapshots
biobouncer_pull("mondo")        # fetch the current MONDO snapshot
biobouncer_cache_dir()          # where snapshots live

remote mode caches responses locally and respects each source's rate limits.

Contributing

Adding a source should be small and declarative: a prefix, a pattern, optional species/version metadata, and (optionally) a cache builder and a remote resolver. See CONTRIBUTING.md and the source-registry spec in PLAN.md.

Acknowledgements

Barret Schloerke and Carson Sievert advise this work as thesis advisors. Posit Software, PBC funded early work on this package and holds copyright together with the author.

License

MIT © biobouncer authors.

Citation

If biobouncer supports your work, please cite it. A preprint is in preparation, and its citation will be added here as soon as it is available. Until then, cite the software archived on Zenodo:

Bharti, S. (2026). biobouncer: Validate Biological Identifiers and Inputs. Zenodo. https://doi.org/10.5281/zenodo.21346522

The DOI above always resolves to the latest release; each release also has its own version DOI on Zenodo. Machine-readable metadata is in CITATION.cff.

About

Best way to validate gene symbols, ontology terms, variant formats, and other biological database IDs in apps and pipelines

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages