UPASHAYA is an R toolkit for investigating how cfDNA fragmentation may contribute to loss of genetic information in liquid biopsy samples. It provides reusable functions and command-line scripts for cohort-level CNV scores, shadow-region analysis, chromatin and gene-group enrichment, and fragment-end motif analysis.
CalculateStoZ.fn.R: calculate and export row-wise Stouffer's Z scores.ShadowRegion.ex.R: filter a cohort table to shadow regions (SZ <= -2).frag2bed.ex.R: convert fragment coordinates to a sorted BED-like table.chrom_enrichment.fn.R: calculate enrichment of shadow regions across chromatin annotations.CNVprofile_plot_data.fn.R: prepare cumulative genomic positions for CNV profile plots.
calculate_motif_frequency_plot.fn.R: create a 5-mer frequency plot from two motif-frequency tables.estimate_shadNshad_ratio.fn.R: calculate shadow/non-shadow motif-frequency ratios.
GeneGroupEnrichment.fn.R: calculate enrichment scores for oncogenes, tumor suppressor genes, human olfactory receptor genes, and housekeeping genes.
R 4.1 or later is recommended. Install the CRAN dependencies with:
install.packages(c("dplyr", "ggplot2", "purrr", "readr", "tidyr", "valr"))Install the Bioconductor dependencies with:
install.packages("BiocManager")
BiocManager::install(c("GenomicRanges", "IRanges", "rtracklayer"))CNV cohort tables must have chr, start, and end as their first three columns, followed by numeric sample columns. Chromosomes are ordered as chr1 through chr22, chrX, and chrY when Stouffer scores are calculated.
Fragment files must be tab-separated and contain genomic chr, start, and end in their first three columns. Additional columns are retained while reading but are not written to the final BED-like output.
Motif tables must contain a 5-mer column in column 1 and numeric motif-frequency columns in the remaining columns. The ratio function expects data frames with mer5 and value columns.
Both scripts accept exactly an input path and an output path:
Rscript CNVShadow/frag2bed.ex.R fragments.tsv fragments.bed
Rscript CNVShadow/ShadowRegion.ex.R cohort.csv shadow_regions.bedfrag2bed.ex.R removes mitochondrial fragments, keeps standard chromosomes, sorts by chromosome and coordinate, and writes four tab-separated columns: chr, start, end, and count.
ShadowRegion.ex.R reads a CSV cohort table, calculates Stouffer's Z scores across the sample columns, retains regions with SZ <= -2, and writes a BED-like TSV containing chr, start, end, and SZ without a header.
Source a function file before calling it:
source("CNVShadow/CalculateStoZ.fn.R")
calculate_stouffers_z("cohort.txt", "cohort_sortedSZ.csv")
source("CNVShadow/CNVprofile_plot_data.fn.R")
plot_data <- CNVprofile_plot_data(cnv_data)
source("FragFront55/calculate_motif_frequency_plot.fn.R")
motif_plot <- calculate_motif_frequency_plot(shadow_start, shadow_end)
library(dplyr)
source("FragFront55/estimate_shadNshad_ratio.fn.R")
motif_ratio <- estimate_shadNshad_ratio(non_shadow_data, shadow_data)chrom_enrichment_score() accepts a cohort BED-like file and a chromatin annotation file:
source("CNVShadow/chrom_enrichment.fn.R")
enrichment <- chrom_enrichment_score(
cohort_file = "Gastric05StoZcnv.bed",
chromatin_file = "withoutCentro.bed"
)
chrom_enrichment_score(
cohort_file = "Gastric05StoZcnv.bed",
chromatin_file = "withoutCentro.bed",
write_csv = TRUE,
out_file = "Gastric05_chromatin_enrichment.csv"
)The cohort file must contain at least five tab-separated columns: chrom, start, end, StoZ, and HScore. The chromatin file must contain chrom, start, end, ChromArm, and ChromStrain. The default threshold is StoZ <= -2, and the default genome length is the hg38 length used by the original analysis: 3,099,734,149 bp.
GeneGroup_enrichmentScore(cohort_name) uses the following files from the current working directory:
TSG_genes.csvONCO_genes.csvhOR_genes.bedHKG_genes.bed<cohort_name>StoZcnv.bed
Example:
source("PenumbraGene/GeneGroupEnrichment.fn.R")
scores <- GeneGroup_enrichmentScore("Gastric05")The cohort file must contain chrom, start, end, StoZ, and HScore. The function returns a named list with enrichment scores for ONCO, TSG, hOR, and HKG. Reference gene files are not included in this repository; record their source, genome build, and version when reproducing an analysis.
This repository contains analysis code only. Cohort data, gene annotations, chromatin annotations, and generated plots or result files are not included. Generated analysis outputs and local editor artifacts are excluded by .gitignore. Run the functions from a working directory containing the required reference files, or update the file paths in the relevant function before execution.
Kabiraj, Debajyoti, Ahmad Salman Sirajee, and Subhajyoti De. "UPASHAYA investigates the consequences of cfDNA fragmentation on the loss of genetic information in liquid biopsy-based cancer detection." npj Genomic Medicine (2026).