A computational pipeline for genotype data processing and GWAS analysis. This repository contains the scripts and utilities required to take raw genotype data through quality control, post-imputation processing, and downstream statistical/survival analysis.
The workflow is divided into logical, independent modules:
utils/: Core data-wrangling scripts (e.g.,prep_vcf.shfor VCF indexing, compression, and tagging, if working with vcf files).post_imputation/: Scripts for handling imputed data chunks (for example: automated logistic regression for specific chromosomes).analysis/: Downstream statistical modeling, including GWAS time-to-event/survival analysis.outputs/: R scripts for generating publication-ready plots (e.g.enhanced_manhattan.Rfor highlighted Manhattan plots).metadata/: Reference tracking and configuration.
To ensure reproducibility, this pipeline relies on Conda.
conda env create -f environment.yml
conda activate gwas_envThis repository contains the computational infrastructure and analytical frameworks only. Due to data privacy and NDAs, all clinical and proprietary genomic data have been omitted.