This repository provides code to compute EVEscape immune escape scores for the M segment proteins of Dabie bandavirus (DBV).
For DBV M segment proteins, EVEscape integrates evolutionary constraints learned from sequence data, structural accessibility inferred from protein structures, and changes in residue physicochemical properties.
The resulting EVEscape scores provide a relative ranking of single amino acid variants based on their potential to escape antibody-mediated immunity while maintaining viral fitness.
Computing EVEscape scores for DBV M segment proteins consists of three components:
-
Fitness
Evolutionary constraint scores derived from an unsupervised generative model trained on multiple sequence alignments of DBV-related viral sequences. -
Accessibility
Structural accessibility estimated from three-dimensional protein structures of DBV M segment proteins, capturing the likelihood that a residue is exposed to antibody binding. -
Dissimilarity
Physicochemical dissimilarity between wild-type and mutant residues, including differences in charge, hydrophobicity, and other residue-level properties that may disrupt antibody–antigen interactions.
The three components are standardized and combined into a single EVEscape score, which is used to rank variants by immune escape potential.
The scripts/ directory contains the code required to compute EVEscape scores for all single amino acid variants of DBV M segment proteins.
-
Step1_train_VAE.sh
Trains the evolutionary model on a multiple sequence alignment of DBV M segment protein sequences. -
Step2_compute_evol_indices_all_singles.sh
Computes evolutionary constraint (fitness) scores for all possible single amino acid substitutions. -
Step3_process_protein_data.py
Processes protein structural data and computes both structural accessibility and physicochemical dissimilarity features. -
Step4_evescape_scores.py
Integrates fitness, accessibility, and dissimilarity components and outputs final EVEscape scores.
To compute EVEscape scores for DBV M segment proteins, the following input data are required:
- One PDB file representing relevant conformations of the DBV M segment proteins
- A multiple sequence alignment (MSA) used to train the evolutionary model
- A FASTA file containing the wild-type DBV M segment protein sequence
The codebase is written in Python and dependencies are managed using Conda.
The environment can be created as follows:
conda env create -f protein_env.yml
conda activate protein_env