Skip to content

Latest commit

 

History

150 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

DelPi

DelPi is an open-source, easy-to-use peptide identification tool for mass spectrometry-based proteomics. It provides a Windows GUI application and a command-line interface, helping users run DIA or DDA peptide searches without building complex workflows by hand.

DelPi applies a pre-trained Transformer encoder to score candidate peptides from raw MS1/MS2 evidence using an acquisition-agnostic representation, enabling a unified workflow across both DIA and DDA data.

DelPi Windows GUI application screenshot

Run peptide identification through an easy-to-use Windows GUI — no manual YAML editing required.

Key Features

  • Deep representation learning: Scores candidate peptides using a pre-trained Transformer encoder, without relying on handcrafted features.
  • DIA and DDA support: Use one workflow across common LC-MS/MS acquisition modes.
  • Library-free search: Generates in silico spectral libraries internally, supporting common PTMs registered in the UniMod database.
  • GPU-accelerated inference: Designed for practical performance on consumer-grade GPUs via PyTorch/CUDA.
  • Experiment-adaptive workflow: Employs a two-stage search with experiment-level transfer learning to adapt to instrument and chromatographic conditions.
  • Easy-to-use Windows GUI: Configure and run DelPi searches without manually editing YAML files.

System Requirements

Memory: ≥ 32 GB RAM

Compute:

  • NVIDIA GPU with CUDA support required
  • Supported OS: Linux, Windows
  • macOS (including Apple Silicon/MPS) and CPU-only execution are not supported

Memory Considerations:

  • DelPi processes input files one run at a time
  • Peak memory usage depends on the size of an individual raw/mzML file being processed
  • Recommended available memory: (single run file size + ~16 GB) to accommodate intermediate data structures, model execution, and OS overhead

Runtime:

  • For a 25 min DIA gradient (human sample, Astral Orbitrap), DelPi completes peptide identification in approximately 15 minutes on a single NVIDIA RTX 4090 GPU

Supported LC-MS/MS Data

  • Formats: mzML and Thermo RAW (other vendor formats can be converted to mzML using ProteoWizard MSConvert)
  • Acquisition modes: DIA and DDA
  • Ion mobility (IM) data (e.g., FAIMS, PASEF) is not currently supported, but will be supported soon

Installation

DelPi is available in two flavors:

  • Windows GUI application — recommended for Windows users who prefer a graphical interface.
  • Command-line tool — cross-platform (Linux/Windows) installation from source.

Option 1: Windows GUI Application

For most Windows users, this is the easiest way to start using DelPi. Download the latest Windows installer (.exe) from the Releases page and run it. The installer bundles all required dependencies; no additional setup is needed.

Option 2: Command-Line Tool

Windows users: Please use PowerShell (not Command Prompt/cmd) for all installation steps. The install scripts (.ps1) require PowerShell to run.

Step 1: Clone the Repository

git clone https://github.com/bertis-informatics/delpi.git
cd delpi

Step 2: Set Up Virtual Environment

Create a virtual environment using venv (Option A) or conda (Option B). Package installation in later steps uses pip or uv.

Option A: Using venv

# Create a virtual environment with Python 3.12
uv venv delpi_env --python 3.12
# If not using uv: python -m venv delpi_env

# Activate the virtual environment
# Windows:
delpi_env\Scripts\activate
# macOS/Linux:
source delpi_env/bin/activate

Option B: Using conda

# Create a conda environment with Python 3.12
conda create -n delpi_env python=3.12 -y

# Activate the environment
conda activate delpi_env

Step 3: Install PyTorch

Visit the PyTorch official website to obtain the appropriate installation command for your system.

Example for CUDA 12.8:

uv pip install torch --index-url https://download.pytorch.org/whl/cu128
# If using pip: pip install torch --index-url https://download.pytorch.org/whl/cu128

Step 4: Install pymsio

pymsio is bundled in the pymsio/ directory. The install script downloads the Thermo RawFileReader DLLs and installs pymsio in one step. For additional details, see the pymsio README.

Windows PowerShell:

.\pymsio\install.ps1

Linux:

chmod +x pymsio/install.sh
./pymsio/install.sh

Step 5: Install DelPi

uv pip install .
# If using pip: pip install .

Step 6: Verify Installation

delpi --help
python -c "import delpi; print('DelPi installed successfully!')"

Quick Test

Verify your DelPi installation using publicly available DIA data from the Skyline tutorial.

Download the Test Data

wget https://skyline.ms/tutorials/DIA-QE.zip
unzip DIA-QE.zip

Quick Start with the Windows GUI

  1. Open the DelPi application.
  2. Select the downloaded raw or mzML input files.
  3. Select a FASTA protein database.
  4. Choose an output folder for search results.
  5. Choose an output spectral library folder.
  6. Review the search settings.
  7. Click Run.
  8. Check the result files in the output folder.

Quick Start with the Command Line

  1. Configure search parameters:

    Copy the example configuration file:

    cp data/example_param.yaml my_config.yaml

    Edit my_config.yaml to specify paths for input_files, fasta_file, output_directory, and database_directory.

  2. Run the search:

    delpi my_config.yaml
  3. Verify output:

    DelPi generates the following files in your specified output_directory:

    • delpi.log: Detailed execution log
    • pmsm_results.<tsv|parquet>: Peptide-spectrum matches with q-values (format depends on configuration; example)
    • protein_group_maxlfq_results.tsv: MaxLFQ protein quantification (example)

    Compare your results with the provided examples to verify correct installation.

Getting Started

1. Prepare LC-MS/MS Data

Ensure your data files are in a supported format. If needed, convert to mzML using ProteoWizard MSConvert.

2. Configure Search Parameters

Create a YAML configuration file based on the example template.

Required fields:

Field Description
acquisition_method Acquisition mode (DIA or DDA)
input_files Paths to LC–MS/MS data files. Accepts a single string or a list, where each entry is either an explicit file path or a glob pattern (*, ?, [], and recursive ** are supported), e.g. /data/*.mzML or /data/**/*.mzML.
fasta_file Protein database in FASTA format
output_directory Directory where search results will be written
database_directory Directory for storing internally generated in silico spectral libraries (if libraries generated using the same FASTA file and search options already exist, they will be reused)

Optional fields:

Digestion and modification parameters can be adjusted for your experimental setup. Modifications can be specified either by their PSI-MS controlled vocabulary names (e.g., Oxidation, Carbamidomethyl) or by their UniMod accession numbers in the UniMod:XX format (e.g., UniMod:35, UniMod:4).

3. Run the Search

Execute DelPi with your configuration file:

delpi /path/to/your/config.yaml

4. Output Files

DelPi generates the following output files:

Main results report, pmsm_results.<tsv|parquet>

Click to expand output fields
Field name Description
frame_num Scan number corresponding to the center of the Peptide–Multi-Spectra Match (PmSM)
run_name Name of the LC–MS run
modified_sequence Peptide sequence including post-translational modifications
precursor_charge Charge state of the precursor ion
sequence_length Length of the peptide sequence
is_decoy Indicator specifying whether the match originates from a decoy sequence
predicted_rt Predicted retention time of the peptide
observed_rt Observed retention time of the peptide
score Raw PmSM score assigned by the DelPi scoring model
global_precursor_q_value Global precursor-level q-value across all runs
global_peptide_q_value Global peptide-level q-value across all runs
global_protein_group_q_value Global protein group-level q-value across all runs
protein_group Protein group inferred according to the parsimony principle (FASTA IDs separated by semicolons)
fasta_id FASTA IDs associated with the peptide, separated by semicolons
precursor_q_value Run-specific precursor-level q-value
peptide_q_value Run-specific peptide-level q-value
protein_group_q_value Run-specific protein group-level q-value
ms1_quantity Integrated area under the precursor ion chromatogram in MS1 spectra
ms2_quantity (DIA only, optional) Precursor abundance quantified from fragment-level signals, before run/RT-dependent normalization
ms2_quantity_normalized (DIA only, optional) ms2_quantity after run/RT-dependent normalization across runs; used as the input to MaxLFQ protein-group quantification

Protein-level quantification results report (DIA only, optional) protein_group_maxlfq_results.<tsv|parquet>

Click to expand output fields
Field name Description
run_name Name of the LC–MS run
protein_group Protein group inferred according to the parsimony principle (FASTA IDs separated by semicolons)
maxlfq_abundance Protein abundance calculated using the MaxLFQ algorithm (Cox et al., 2014) from normalized precursor quantities (ms2_quantity_normalized)

Citation

If you use DelPi in your research, please cite:

Park, J., Kim, K., Kang, U.-B., & Kim, S. DelPi Learns Generalizable Peptide–Signal Correspondence for Mass Spectrometry-Based Proteomics. bioRxiv (2026). https://doi.org/10.64898/2026.01.06.697814

License

DelPi is freely available under the MIT License.

Contact

For questions, bug reports, or feature requests, please contact Jungkap Park, Ph.D. at jungkap.park@bertis.com

About

DelPi: Deep Learning-based Peptide Identification Search Engine

Topics

Resources

Stars

11 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages