Skip to content
 
 

Repository files navigation

ProDomino

ProDomino prioritizes positions in a host protein for experimental domain-insertion engineering. It combines residue-level embeddings from the ESM-2 3B protein language model with a small neural-network decoder and returns one insertion-tolerance score per host residue.

The associated paper is Rational engineering of allosteric protein switches by in silico prediction of domain insertion sites.

Important

ProDomino predicts generic host-site tolerance. It does not predict whether a particular inserted domain will produce an allosteric switch, and its scores are ranking values rather than calibrated probabilities. Read the model card before designing experiments.

Installation

The reference environment uses Python 3.10, PyTorch 2.1, and CUDA 12.1. A CUDA GPU with substantial memory is recommended for the 3-billion-parameter ESM model. environment.yml is specifically the CUDA 12.1 environment; for CPU-only or macOS use, install the platform-appropriate PyTorch 2.1 build first and then install this package with pip.

conda env create --file environment.yml
conda activate prodomino

The environment installs this repository in editable mode. If you already have compatible dependencies, install the package directly:

python -m pip install -e . --no-deps

To add the practical-notebook and test tooling to an existing Python 3.10/3.11 environment:

python -m pip install -e ".[notebook,test]"

CPU inference is supported but is slow and memory intensive. Embeddings can instead be generated once on a GPU, saved as NumPy arrays, and scored later on a CPU.

The first Embedder initialization downloads the multi-gigabyte ESM-2 weights into the Torch Hub cache. On an offline cluster, pre-populate a persistent cache on a connected machine and point TORCH_HOME to it. Ensure that the job has enough cache space and system/GPU memory before loading the model.

The decoder loader accepts only bytes matching the released checkpoint's allowlisted SHA-256 digest. This is deliberate: the pinned research environment's PyTorch loader is not a security sandbox, even in weights_only mode. The fair-esm model loader also uses PyTorch serialization internally, so only use official ESM weights from a trusted download/cache. Custom decoder weights require a separately reviewed, non-pickle loading path.

Quick start

from pathlib import Path

from Bio import SeqIO
import torch
from ProDomino import Embedder, ProDomino

record = next(
    record
    for record in SeqIO.parse("example_sequences.fasta", "fasta")
    if record.id == "AraC"
)
sequence = str(record.seq)
checkpoint = Path("checkpoints/main_checkpoint.ckpt")
device = "cuda" if torch.cuda.is_available() else "cpu"

embedder = Embedder(device=device)
model = ProDomino(checkpoint, model_type="mini_3b_mlp", device=device)

embedding = embedder.predict_embedding(sequence)
prediction = model.predict_insertion_sites(
    embedding,
    sequence=sequence,
    embedding_model_name=embedder.model_name,
    embedding_layer=embedder.repr_layer,
)

prediction.show_trace(show_top_hits=True)
candidates = prediction.candidate_table(
    n=12,
    min_spacing=5,
    one_based=True,
)
for candidate in candidates:
    print(candidate)

predict_sequence(sequence, embedder) supplies embedding provenance automatically. When scoring a cached NumPy array as above, embedding_model_name and embedding_layer are required so a same-width but incompatible representation is not accepted silently. The practical notebook stores and validates these fields beside every cached array.

Candidate positions use an explicit insertion-after convention: one-based position 42 means insert between host residues 42 and 43. Set one_based=False only when array indices are required.

For a complete design-to-export workflow, open practical_design.ipynb. It includes embedding caching, score clustering, structural review, interface exclusions, controls, and CSV/JSON export. The older example.ipynb is retained as a historical paper example and contains machine-specific paths; do not use Run All without reviewing it.

What the model computes

amino-acid sequence
        │
        ▼
ESM-2 3B, layer 36 residue embeddings (L × 2560)
        │
        ▼
Linear(2560, 1280) → ReLU → Linear(1280, 1) → sigmoid
        │
        ▼
one ranking score per host residue (L)

Fair-esm commonly uses a 1,022-residue extraction length for this model. ESM-2's rotary positions permit longer inputs, so the embedder warns and continues without silently truncating or stitching. Treat results beyond that reference length cautiously: they are less validated, and attention memory grows approximately quadratically with sequence length.

Practical interpretation

  • Rank regions, not just isolated maxima. Neighboring high-scoring residues often describe one tolerant loop.
  • Enforce spacing between candidates so a short list samples distinct structural regions.
  • Inspect solvent exposure, secondary structure, active sites, ligand contacts, interfaces, and known functional motifs before ordering constructs.
  • Include high-, moderate-, and low-scoring controls. The comparison is more informative than testing only predicted winners.
  • Record the exact host sequence, residue-numbering convention, inserted domain, linker sequences, and construct junctions.
  • Treat predictions as hypotheses requiring experimental validation.

Structure annotation

InsertionSitePrediction.add_pdb_file(...) aligns the selected PDB chain sequence to the scored host sequence before transferring scores. This avoids relying on raw PDB residue numbers, which can contain gaps and insertion codes. Inspect alignment warnings, especially for engineered structures or partial chains.

The annotated structure stores mapped prediction scores from 0–100 in the B-factor field and uses -1 for unmapped residues in the selected chain. That operation overwrites crystallographic B factors in the in-memory copy, so keep the original PDB file unchanged.

Repository scope and reproducibility

This repository contains the inference implementation, released checkpoint, examples, and selected supplementary data. It does not contain the original training dataset construction pipeline, data split definitions, full training command/configuration, or a complete paper-reproduction workflow. Consequently, the released predictor can be run and inspected, but the published checkpoint cannot be independently regenerated from this repository alone. Those limits are documented in the model card.

Development

Run the test suite from the repository root:

pytest

License and citation

The software is released under CC BY-NC 4.0. Cite the paper using CITATION.cff. The license of the published article and third-party models/data may differ and must be considered separately.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages