Skip to content
View clintval's full-sized avatar

Block or report clintval

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
clintval/README.md

Cover

I lead technical teams in biotech and write software for new genomics technologies. At Fulcrum Genomics you'll find me building tools and leading others in the fields of oncology, cell & gene editing, and precision medicine all while ensuring we deliver high-quality services to our clients and partners.

Featured

ProjectStackInstallWhat it does
chum Language Install with bioconda Evaluate baits in a hybrid selection panel.
krak Language Install with bioconda An addicting set of Kraken-enhancing tools.
unmux Language Install with bioconda Parse and demultiplex records, splitcode-style.
vartovcf Language Install with bioconda Stream VarDict variants into VCF v4.2.
neodisambiguate Language Install with bioconda Disambiguate reads mapped to multiple references.
bedspec Language PyPI Release An HTS-specs compliant BED toolkit.
cellme Language PyPI Release Make a truth-track VCF of a cell line's known mutations.
pybgzf Language PyPI Release Streaming BGZF with on-the-fly tabix and CSI indexing.
typeline Language PyPI Release Dataclasses to delimited text, round-trip with types.

chum

Score capture baits against a reference:

❯ chum score \
    --baits baits.fa \
    --targets targets.bed \
    --reference hg38.fa \
    --per-bait per-bait.tsv

krak

Bridge Kraken classifications into a BAM and filter by taxon:

❯ krak annotate \
      -i input.bam \
      -d /kraken-db \
      -a <(krak prep input.bam | kraken2 --db /kraken-db --output - -) \
  | krak filter -t 9606 -o output.bam

unmux

Demultiplex a dual-index paired-end run against a sample sheet, routing each read pair by its i7+i5 barcode concatenation:

❯ unmux "R1.fastq.gz" "I1.fastq.gz" "I2.fastq.gz" "R2.fastq.gz" \
    --extract "i7=1:0:8" \
    --extract "i5=2:0:8" \
    --extract "r1=0:0:end" \
    --extract "r2=3:0:end" \
    --group "samples=metadata.tsv" \
    --group "samples::match=i7+i5" \
    --template "r1" \
    --template "r2" \
    --sample-from-group "samples" \
    --out "demux/%sample.R%ordinal.fq"
    

neodisambiguate

Disambiguate templates aligned to human and mouse references:

❯ neodisambiguate \
    --input dna00001.aligned-to-human.bam dna00001.aligned-to-mouse.bam \
    --output out/dna00001 \
    --names hg38 mm10

cellme

Write an hg38 truth track of the known mutations in the MOLT-4 cell line:

❯ cellme "MOLT-4" \
    --build hg38 \
    --reference hg38.fa \
    --output MOLT-4.hg38.vcf.gz

pybgzf

Write BED lines as BGZF, indexed as they are written, then query a region:

import pybgzf
from pybgzf import Columns, IndexFormat

with pybgzf.writer("features.bed.gz", index=IndexFormat.TBI, columns=Columns.BED) as handle:
    handle.write("chr1\t100\t200\tgene-a\n")
    handle.write("chr1\t150\t300\tgene-b\n")

with pybgzf.IndexedReader("features.bed.gz") as reader:
    for line in reader.query("chr1", 180, 190):
        print(line)

Elsewhere

LinkedIn Fulcrum Genomics Bioconda PyPI

Pinned Loading

  1. unmux unmux Public

    Flexible read parsing and demultiplexing to FASTX/SAM/BAM/CRAM

    Rust 2

  2. pybgzf pybgzf Public

    Streaming BGZF compression with on-the-fly tabix and CSI indexing

    Rust 1

  3. bedspec bedspec Public

    An HTS-specs compliant BED toolkit

    Python 3 1

  4. typeline typeline Public

    Write dataclasses to delimited text formats and read them back again

    Python

  5. neodisambiguate neodisambiguate Public

    Disambiguate reads that were mapped to multiple references

    Scala 4 1

  6. fg-labs/chum fg-labs/chum Public

    Evaluate the effectiveness of baits in a hybrid selection panel

    Rust 4