This repository serves two purposes: the original Genomics for Software Engineers course in
Course/, and the Genomics .NET platform it grew into. Course returners → start here.
Genomics is a .NET-native platform for genomic data processing: a strongly typed domain model, streaming parsers for the major sequencing file formats, sequence operations, indexing, algorithms, and composable pipelines: explicit interop with the established bioinformatics ecosystem (samtools, bcftools, minimap2, BWA, GATK) instead of reinventing it.
For .NET engineers, genomic data should feel like a first-class domain, not something you must always hand off to Python, Java, or C++ pipelines. That is what this platform is for.
If you came here for the Genomics for Software Engineers curriculum, it lives on in Course/: the 12-week, hands-on crash course that teaches minimum-viable biology and maximum practical bioinformatics through software analogies, with a printable single-document version (genomics.docx).
The course is more than documentation: it is the learning foundation of this platform. Every concept taught there (coordinates, sequences, file formats, reads, alignments, variants) is being turned into a typed, tested, benchmarked primitive in the codebase. Learn the field in Course/, then follow how it becomes code here.
This repository grew out of a software engineer's ramp-up: joining a genomics project, finding no material between "PhD-level biology" and "tool docs for people who already know the field," and building a self-study curriculum from first principles: DNA as a data structure, the sequencer as a physical-to-digital converter, alignment as string matching, FASTA and VCF as data contracts.
| Path | Purpose |
|---|---|
Course/ |
The Genomics for Software Engineers curriculum + printable genomics.docx |
docs/guides/ |
Step-by-step guides: quick start, then each file format (FASTQ, SAM, BAM, BAM indexes, VCF) and the sequence algorithms |
Genomics.Core/ |
Domain foundation: 0-based/1-based genomic coordinates, chromosomes and contigs, loci and regions with interval algebra (distance, intersect, merge, subtract), normalized interval/region sets (union, intersect, subtract, complement) and a settled IntervalTree overlap index — with validation and the shared GenomicsException family every package builds on |
Genomics.Sequences/ |
Typed molecular sequences over a compact Nucleotide primitive: DnaSequence, RnaSequence, ProteinSequence with strict alphabet validation, reverse complement, transcription, translation via the standard genetic code, and content-based equality |
Genomics.Sequencing/ |
The sequencing domain: typed Read objects over DnaSequence + Phred QualityScores, and ReadPair groupings: the model FASTQ files load into |
Genomics.IO/ |
Streaming, async, cancellable file-format readers and writers: FASTA, FASTQ (plain or gzip, auto-detected), VCF (headers, typed INFO/FORMAT, genotypes, gzip input), SAM (typed CIGAR/flags/tags, Alignment, gzip input), and BAM (BGZF block I/O, the binary header and record codecs, SAM⇄BAM round trips, BAI index build/.bai sidecar I/O, virtual-offset random access, and indexed region queries) |
Genomics.Console/ |
A self-contained hands-on tour of the whole platform: run the tour |
Genomics.Tests/ |
Unit tests: one project, folders per namespace (Core/, Sequences/, IO/, …) |
Genomics.Benchmarks/ |
BenchmarkDotNet suites: throughput, allocations, and scaling curves for each foundation layer as features ship |
Want to see the platform work end to end without writing any code? The
Genomics.Console app walks through coordinates, sequence algorithms, and
every supported file format using synthetic data it generates itself:
dotnet run --project Genomics.ConsoleIt is a sample application (never packed or published), and verify runs
every scenario as a self-test. See
the tour guide for commands and options.
The platform — all code, tests, benchmarks, and project documentation — is licensed under Apache-2.0. The course materials in Course/ are licensed under MIT.