Skip to content

Latest commit

 

History

26 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

GenAnalyzer – C++ SNP Risk Analysis

GenAnalyzer is an object-oriented C++ project for analysing raw genetic data, for example the text files exported by AncestryDNA.

Consumer DNA tests such as AncestryDNA genotype SNPs (single nucleotide polymorphisms), which are single-base variants in the genome, and give the raw data back to the person tested. In principle these SNPs can be used to spot genetic predispositions for diseases or to inform lifestyle decisions (e.g. diet, micronutrients, detoxification pathways).

GenAnalyzer automates this comparison: it reads a personal SNP profile and compares it against a predefined list of risk variants (e.g. for MCAS, methylation disorders, cancer). Notable genotypes are reported together with the affected gene's function and a simple risk estimate.


Structure & how it works

  • Input: AncestryDNA raw data (.txt) and risk-SNP tables (.tsv)
  • Analysis: every SNP in the genome is compared with the known risk variants
  • Risk scoring: simple point system (1 point = heterozygous, 2 points = homozygous)
  • Output: terminal summary and text export

Example output

The program's output is in German:

Risiko-Score: 5 → Mäßig erhöht
rs1801133   AG   Heterozygot   MTHFR   Methylierung (C677T)
rs4680      AA   Homozygot     COMT    Dopaminabbau (Val/Met)

(Risiko-Score = risk score, Mäßig erhöht = moderately elevated, Heterozygot/Homozygot = heterozygous/homozygous.)


Project structure

GenAnalyzer/
├── src/                # main.cpp
├── lib/                # implementations (SNP, Genome, Analyzer, Disease)
├── include/            # header files
├── data/               # sample SNP data (MCAS_snps.tsv etc.)
├── build/              # (generated by CMake)
├── CMakeLists.txt
└── README.md

Note on the risk score

The scoring used in this project is a simplified heuristic for demonstration purposes only. It is not a medical risk assessment.

It is loosely based on:

Study: "Population-standardized genetic risk score"
GRS-RAC model (Genetic Risk Score – Risk Allele Count)

RR = 2 points → homozygous risk (two risk alleles) RN = 1 point → heterozygous (one risk allele, one normal allele) NN = 0 points → homozygous normal (no risk alleles)

The sum of all points is the individual risk score, which this project groups into three classes: low (Gering), moderately elevated (Mäßig erhöht) and high (Hoch). The distribution of risk alleles, their population frequency and interactions with other genes are not taken into account.


Build & run

mkdir build
cd build
cmake ..
make
cd ..
.\build\GenAnalyzer

Note: If the program is started from inside the build/ directory, it cannot find the files in the data/ folder.

Fix:

  • Go up one level and start GenAnalyzer from there -> .\build\GenAnalyzer
  • This keeps the working directory correct, so the relative paths to data/ work as intended.

Dependencies

  • C++17
  • CMake ≥ 3.10
  • MSYS2 / GCC or Visual Studio Code with CMake Tools

Features

GenAnalyzer's main functionality is split into modular classes:

  • Genome: loads and stores the SNP data of a raw genetic dataset
  • Disease: holds the risk SNPs for one disease
  • Analyzer: compares a genome with the risk SNPs and scores the genetic risk

main() runs an example analysis, in which you can:

  • load a genome,
  • select one or more diseases,
  • run the analysis,
  • view the results and
  • export a results report.

After a successful analysis a file is created automatically in data/output/, e.g.: data/output/DemoSample_results.txt

This file serves as a demo of the analysis output.


Adding new diseases

GenAnalyzer automatically picks up every .tsv disease file in data/disease/. No code changes are needed.

  1. Create a file in data/disease/, e.g.:
Type2Diabetes.tsv
  1. Add the following structure:
rsID        gene        function
rs1801282   PPARG       Insulin sensitivity / adipogenesis
rs7754840   CDKAL1      Insulin secretion / beta cells
rs13266634  SLC30A8     Zinc transporter / glucose homeostasis
rs5219      KCNJ11      Potassium channel / insulin release

🔹 Note: columns must be separated by tabs, not commas or spaces!

  1. Restart the program
    • The file is detected automatically
    • "Type2Diabetes" appears in the selection menu

Possible extensions

  • CSV/HTML output
  • Extend the genome model with personal data (age, BMI, lifestyle)
  • RiskAnalyzer v2: weighted risk alleles, better visualisation
  • Terminal UI with menu navigation

Author: Niklas Mitterbuchner Project for: C++ software development course, final project (summer semester)

About

Object-oriented C++ tool that matches AncestryDNA raw data against risk-SNP lists (e.g. MTHFR, COMT, MCAS) and outputs a simple risk score.

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages