Skip to content

Latest commit

 

History

17 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🧬 GWAS QC & Downstream Analysis Pipeline

Bash R Perl

A computational pipeline for genotype data processing and GWAS analysis. This repository contains the scripts and utilities required to take raw genotype data through quality control, post-imputation processing, and downstream statistical/survival analysis.

📂 Repository Structure

The workflow is divided into logical, independent modules:

  • utils/: Core data-wrangling scripts (e.g., prep_vcf.sh for VCF indexing, compression, and tagging, if working with vcf files).
  • post_imputation/: Scripts for handling imputed data chunks (for example: automated logistic regression for specific chromosomes).
  • analysis/: Downstream statistical modeling, including GWAS time-to-event/survival analysis.
  • outputs/: R scripts for generating publication-ready plots (e.g. enhanced_manhattan.R for highlighted Manhattan plots).
  • metadata/: Reference tracking and configuration.

Environment Setup

To ensure reproducibility, this pipeline relies on Conda.

conda env create -f environment.yml
conda activate gwas_env

⚠️ Note on Data Privacy

This repository contains the computational infrastructure and analytical frameworks only. Due to data privacy and NDAs, all clinical and proprietary genomic data have been omitted.

About

A scalable pipeline for GWAS quality control, imputation processing, and survival analysis.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages