Skip to content

Repository files navigation

biocohort biocohort hex logo

CRAN status Lifecycle: experimental R-CMD-check r-universe DOI License: MIT

biocohort keeps the subjects, samples, and analysis outputs of a study in one validated object. Species and assays are values in the data, not columns or classes, so the same functions work for any organism and any omics assay.

Documentation: https://www.samuelbharti.com/biocohort/

Installation

From CRAN:

install.packages("biocohort")

The development version, from GitHub:

pak::pak("samuelbharti/biocohort/pkg-r")

or from r-universe:

install.packages("biocohort", repos = "https://samuelbharti.r-universe.dev")

Usage

A manifest is one long-format table, one row per sample. Four columns carry the shape of the study: subject_id, assay, sample_id, role. Everything else is metadata.

subject_id,species,genotype,assay,sample_id,role
R1,rat,WT,wes,T1,tumor
R1,rat,WT,wes,N1,normal
R2,rat,KO,wes,T2,tumor
R2,rat,KO,wes,N2,normal
library(biocohort)

parsed <- read_manifest("manifest.csv")
cohort <- cohort_new(parsed$subject_tbl, parsed$sample_map)

pkg-r/README.md carries the rest, including the sample sheet a pipeline reads and the record of every manual fix.

Motivation

Every study starts the same way. A spreadsheet of subjects, a folder of sample IDs that do not quite match it, and a script that fixes the mismatch by hand and then gets rewritten for the next study. biocohort solves that once instead of once per study: one manifest becomes one checked object, it writes the sample sheet the pipeline wants, and it records every manual fix so the reason survives six months.

Species and assay are plain values in the data rather than something the code checks against a list, so rat, mouse or human all work, and so does whatever the lab runs that week.

Bioconductor already has MultiAssayExperiment for lining up data that is already loaded. biocohort sits one step earlier, before anything is loaded: the manifest, the file paths, the sample sheet, the record of every fix. It is the paperwork step before MultiAssayExperiment rather than a replacement for it.

Repository layout

The R package lives in pkg-r/, not at the repository root, so package commands run against that path. See CONTRIBUTING.md for the workflow.

Documentation

pkg-r/README.md has a runnable quick start and the full list of what the package does. The docs site has three articles:

Citation

The package is archived on Zenodo. Use the concept DOI, which always resolves to the newest archived release:

Bharti, S. (2026). biocohort: Cohort Objects for Subjects and Samples in Omics Studies. Zenodo. https://doi.org/10.5281/zenodo.22685057

To pin the exact version you used, take its DOI from CITATION.cff, which carries one identifier per archived release.

In R, citation("biocohort") prints the same reference, and CITATION.cff carries the same metadata for the "Cite this repository" button on GitHub.

Contributing

See CONTRIBUTING.md for the development workflow.

License

MIT. See LICENSE.

About

Keeps the subjects, samples, and analysis outputs of a genomics (or any omics) study in one validated R object.

Resources

Code of conduct

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages