nf-core/msproteomics is a mass spectrometry proteomics preprocessing pipeline supporting DIA, DDA LFQ, TMT label check, and generic FragPipe workflows. This document describes how to prepare input data, configure the pipeline, and run each workflow type.
New to the pipeline? Start with the Quick Start Guide for a 5-minute introduction. For database preparation, see the Database Preparation Guide. For pipeline design details, see the Workflow Architecture.
The pipeline accepts a CSV samplesheet describing your samples and their raw data file locations.
--input samplesheet.csvThe samplesheet must contain the following columns:
| Column | Description | Required |
|---|---|---|
sample |
Unique sample identifier | Yes |
spectra |
Path or URI to the raw/mzML data file | Yes |
condition |
Experimental condition or group (used for downstream statistics; defaults to sample name) | No |
label |
Isobaric label channel (e.g., TMT126, TMT127N); leave empty for LFQ/DIA | No |
fraction |
Fraction number for fractionated experiments; leave empty for single-shot | No |
Example samplesheet (DIA or DDA LFQ):
sample,spectra,condition,label,fraction
Sample_A1,/data/raw/A1.raw,control,,
Sample_A2,/data/raw/A2.raw,control,,
Sample_B1,/data/raw/B1.raw,treated,,
Sample_B2,/data/raw/B2.raw,treated,,Example samplesheet (TMT):
sample,spectra,condition,label,fraction
Sample_1,/data/raw/pool1.raw,control,TMT126,
Sample_1,/data/raw/pool1.raw,control,TMT127N,
Sample_1,/data/raw/pool1.raw,treated,TMT128C,The pipeline internally generates an SDRF file from the samplesheet via the GENERATE_SDRF_FROM_SAMPLESHEET module.
This SDRF is used for bookkeeping and compatibility with downstream tools that expect SDRF-formatted metadata.
The --mode parameter selects the analysis engine:
| Mode | Description | Engine |
|---|---|---|
diann |
DIA quantitative proteomics | DIA-NN |
fragpipe |
DDA LFQ, TMT, or generic FragPipe workflows | FragPipe |
DIA workflows use DIA-NN for library-free or library-based DIA analysis.
DIA-specific variants (e.g., phospho) are applied via -c conf/variants/*.config files (see DIA Method Variants).
nextflow run nf-core/msproteomics \
--mode diann \
--input samplesheet.csv \
--database /path/to/database.fasta \
--outdir results \
-profile dockerIf --database is omitted, the UniProt reference proteome for --organism (default Homo sapiens) is downloaded at runtime from conf/reference_proteomes.config; see Database Preparation.
--database /path/to/database.fasta always takes precedence.
FragPipe mode supports DDA LFQ, TMT label check, TMT quantification, and generic FragPipe workflows.
The specific workflow is controlled by the --fragpipe_workflow and --tmt_mode parameters.
DDA LFQ -- uses MSFragger, MSBooster, Percolator, ProteinProphet, Philosopher, and IonQuant:
nextflow run nf-core/msproteomics \
--mode fragpipe \
--fragpipe_container your-registry/fragpipe:24.0 \
--fragpipe_workflow LFQ-MBR.workflow \
--input samplesheet.csv \
--database /path/to/uniprot_human.fasta \
--outdir results \
-profile dockerTMT Label Check -- assess TMT labeling efficiency:
nextflow run nf-core/msproteomics \
--mode fragpipe \
--fragpipe_container your-registry/fragpipe:24.0 \
--tmt_mode labelcheck \
--input samplesheet.csv \
--database /path/to/uniprot_human.fasta \
--outdir results \
-profile dockerTMT Quantification -- full TMT isobaric quantification:
nextflow run nf-core/msproteomics \
--mode fragpipe \
--fragpipe_container your-registry/fragpipe:24.0 \
--tmt_mode quant \
--input samplesheet.csv \
--database /path/to/uniprot_human.fasta \
--outdir results \
-profile dockerGeneric FragPipe -- run any FragPipe workflow by providing a .workflow configuration file:
nextflow run nf-core/msproteomics \
--mode fragpipe \
--fragpipe_container your-registry/fragpipe:24.0 \
--fragpipe_workflow /path/to/custom.workflow \
--input samplesheet.csv \
--database /path/to/uniprot_human.fasta \
--outdir results \
-profile dockerFragPipe .workflow files are key=value configuration files that control which tools run and their parameters.
The pipeline ships three built-in workflow files in assets/:
| File | Use Case | Key Tools |
|---|---|---|
LFQ-MBR.workflow |
DDA label-free quantification with match-between-runs | MSFragger, MSBooster, Percolator, ProteinProphet, IonQuant |
TMT-labelcheck-229.workflow |
TMT label check for TMT6/10/11 (mass 229 Da) | MSFragger, Percolator, ProteinProphet |
TMT-labelcheck-304.workflow |
TMT label check for TMTpro/16/18 (mass 304 Da) | MSFragger, Percolator, ProteinProphet |
To create a custom .workflow file:
- Open the FragPipe GUI and configure your analysis
- Save the workflow configuration (File > Export Workflow)
- Provide the exported file to the pipeline with
--fragpipe_workflow /path/to/custom.workflow
Key parameters users commonly tune in .workflow files:
- Enzyme: trypsin, lysc, chymotrypsin, etc.
- Missed cleavages: typically 1-2
- Variable modifications: oxidation (M), acetylation (protein N-term)
- Mass tolerances: precursor and fragment mass tolerances
- Calibrate mass: 0 (disabled), 1 (single pass), 2 (two pass, recommended)
Note
TMT label check workflows use TMT as a variable modification (not fixed) to detect unlabeled peptides. This is intentional and differs from standard TMT quantification workflows.
The --fragpipe_mode parameter controls execution mode:
pipeline(default): Modular Nextflow subworkflow executionheadless: Single-process FragPipe headless execution
| Parameter | Description | Default |
|---|---|---|
--input |
Path to CSV samplesheet | Required |
--mode |
Analysis mode: diann or fragpipe |
Required |
--outdir |
Output directory | Required |
--database |
FASTA protein database | UniProt reference proteome for --organism |
--fragpipe_workflow |
FragPipe .workflow file or workflow name (FragPipe mode) |
None |
--tmt_mode |
TMT analysis mode: labelcheck or quant (FragPipe mode) |
None |
--fragpipe_mode |
FragPipe execution mode: pipeline or headless |
pipeline |
DIA workflows use publicly available containers for DIA-NN and associated tools. No special licensing is required.
The pipeline defaults to the public base image docker.io/fcyucn/fragpipe:24.0, which includes open-source FragPipe tools (Philosopher, Percolator, MSBooster, PTMShepherd, etc.).
However, commercial/licensed tools (MSFragger, IonQuant, DiaTracer) are not included in the public image. You must build your own container image with these tools under your own license terms. See Docker build instructions for academic and commercial container builds.
FragPipe is free for academic use. Commercial users must obtain a license from Fragmatics.
Use -profile to select a software packaging method:
| Profile | Description |
|---|---|
docker |
Run with Docker containers (recommended) |
singularity |
Run with Singularity containers |
apptainer |
Run with Apptainer containers |
podman |
Run with Podman containers |
conda |
Run with Conda environments (last resort) |
test |
Minimal stub test for CI validation (DIA mode) |
test_dia |
DIA workflow test with small DIA dataset |
test_dda_lfq |
DDA LFQ workflow test with small DDA dataset |
test_tmt |
TMT label check workflow test |
test_tmtq |
TMT quantification workflow test |
Multiple profiles can be combined: -profile test_dia,docker
nextflow run nf-core/msproteomics \
--mode diann \
--input samplesheet.csv \
--database /path/to/database.fasta \
--outdir results \
-profile dockernextflow run nf-core/msproteomics -profile docker -params-file params.yamlinput: "samplesheet.csv"
mode: "fragpipe"
fragpipe_container: "your-registry/fragpipe:24.0"
fragpipe_workflow: "LFQ-MBR.workflow"
database: "/path/to/uniprot_human.fasta"
outdir: "results"nextflow run nf-core/msproteomics \
--mode diann \
--input samplesheet.csv \
--database /path/to/database.fasta \
--outdir results \
-profile docker \
-resumenextflow pull nf-core/msproteomicsSpecify a pipeline version with -r to ensure reproducible results:
nextflow run nf-core/msproteomics -r 1.0.0 \
--mode diann \
--input samplesheet.csv \
--database /path/to/database.fasta \
--outdir results \
-profile dockerThe pipeline loads configuration files automatically based on the --mode parameter:
DIA workflows (--mode diann):
nextflow.config-- root config, always loadedconf/base_configs/diann.config-- DIA-NN computed params and helper functionsconf/variants/*.config-- variant-specific overrides (e.g., diaphos) applied via-cconf/instruments/*.config-- instrument-specific overrides applied via-c
FragPipe workflows (--mode fragpipe):
nextflow.config-- root config, always loaded (all params defined here)--tmt_typeselects the TMT plex (TMT6, TMT10, TMT11, TMT16, TMT18, TMTPRO); the correct labelcheck workflow file is auto-selectedconf/instruments/*.config-- instrument-specific overrides applied via-c
To change compute resources for specific processes, see the nf-core resource tuning documentation.
To pass additional arguments to specific tools, use the ext.args directive in a custom config file.
See the nf-core tool arguments documentation.
Resource requirements depend on dataset size and analysis type:
| Dataset Size | Files | Recommended Memory | Notes |
|---|---|---|---|
| Small | < 20 | 16 GB | Default resources sufficient |
| Medium | 20-100 | 32 GB | Increase MSFragger and IonQuant memory |
| Large | 100+ | 64 GB | Consider splitting into batches |
- DIA-NN library generation is the most memory-intensive step; allocate 32-64 GB for large proteomes or phosphoproteomics
- MSFragger memory scales with database size; large databases (>100,000 sequences) may need 32+ GB
- FragPipe headless mode runs all tools in a single process and requires memory sufficient for the largest tool
- FragPipe pipeline mode distributes work across processes, allowing smaller per-process memory allocation
For cloud environments, memory-optimized instances are recommended (e.g., AWS r6i family, GCP n2-highmem).
Warning
Do not use -c <file> to specify pipeline parameters.
Custom config files specified with -c must only be used for resource tuning, output directories, or module arguments (ext.args).
The pipeline dynamically loads configurations from nf-core/configs. Check if your institution has a pre-configured profile available.
Set the Nextflow version in your launch pre-run script:
export NXF_VER=25.10.2FragPipe processes require compute environments with sufficient memory.
Recommended instance families: r6i (memory-optimized).
DIA-NN in-silico library generation uses different m/z ranges and mass accuracy settings depending on the mass spectrometer.
The pipeline ships per-instrument config files in conf/instruments/ that override the defaults when included with -c.
| Abbreviation | Instrument | Config File | diann_min_pr_mz | diann_max_pr_mz | diann_min_fr_mz | diann_max_fr_mz | diann_library_mass_acc | diann_library_ms1_acc |
|---|---|---|---|---|---|---|---|---|
| ASC | Thermo Ascend | conf/instruments/thermo_ascend.config (default) |
350 | 1050 | 200 | 1800 | 18 | 5 |
| FLX | Bruker timsTOF Flex | conf/instruments/bruker_flex.config |
100 | 1700 | 100 | 1700 | 15 | 15 |
Usage example:
nextflow run nf-core/msproteomics \
--mode diann \
--input samplesheet.csv \
--database /path/to/database.fasta \
--outdir results \
-c conf/instruments/bruker_flex.config \
-profile dockerThermo Ascend users can omit the -c flag since the defaults are already optimized for Ascend instruments.
Specialized DIA analysis modes are available as variant config files in conf/variants/.
These configs adjust DIA-NN parameters (e.g., variable modifications, phosphosite monitoring) for specific experimental designs.
| Variant | Config File | Description |
|---|---|---|
| DIA Phosphoproteomics | conf/variants/diaphos.config |
Phosphosite localization with UniMod:21 monitoring |
Usage example:
nextflow run nf-core/msproteomics \
--mode diann \
--input samplesheet.csv \
--database /path/to/database.fasta \
--outdir results \
-c conf/variants/diaphos.config \
-profile dockerVariant configs can be combined with instrument configs by specifying multiple -c flags:
nextflow run nf-core/msproteomics \
--mode diann \
--input samplesheet.csv \
--database /path/to/database.fasta \
--outdir results \
-c conf/instruments/bruker_flex.config \
-c conf/variants/diaphos.config \
-profile dockerTo limit Nextflow's JVM memory usage, add the following to your environment:
NXF_OPTS='-Xms1g -Xmx4g'- Quick Start Guide -- Get running in 5 minutes
- Output Documentation -- Interpreting pipeline results
- Database Preparation -- Choosing and preparing FASTA databases
- Workflow Architecture -- Pipeline design and workflow routing
- Troubleshooting -- Common errors and solutions