Skip to content

Latest commit

 

History

History
420 lines (310 loc) · 17 KB

File metadata and controls

420 lines (310 loc) · 17 KB

nf-core/msproteomics: Usage

Introduction

nf-core/msproteomics is a mass spectrometry proteomics preprocessing pipeline supporting DIA, DDA LFQ, TMT label check, and generic FragPipe workflows. This document describes how to prepare input data, configure the pipeline, and run each workflow type.

New to the pipeline? Start with the Quick Start Guide for a 5-minute introduction. For database preparation, see the Database Preparation Guide. For pipeline design details, see the Workflow Architecture.

Input Format

CSV Samplesheet

The pipeline accepts a CSV samplesheet describing your samples and their raw data file locations.

--input samplesheet.csv

The samplesheet must contain the following columns:

Column Description Required
sample Unique sample identifier Yes
spectra Path or URI to the raw/mzML data file Yes
condition Experimental condition or group (used for downstream statistics; defaults to sample name) No
label Isobaric label channel (e.g., TMT126, TMT127N); leave empty for LFQ/DIA No
fraction Fraction number for fractionated experiments; leave empty for single-shot No

Example samplesheet (DIA or DDA LFQ):

sample,spectra,condition,label,fraction
Sample_A1,/data/raw/A1.raw,control,,
Sample_A2,/data/raw/A2.raw,control,,
Sample_B1,/data/raw/B1.raw,treated,,
Sample_B2,/data/raw/B2.raw,treated,,

Example samplesheet (TMT):

sample,spectra,condition,label,fraction
Sample_1,/data/raw/pool1.raw,control,TMT126,
Sample_1,/data/raw/pool1.raw,control,TMT127N,
Sample_1,/data/raw/pool1.raw,treated,TMT128C,

The pipeline internally generates an SDRF file from the samplesheet via the GENERATE_SDRF_FROM_SAMPLESHEET module. This SDRF is used for bookkeeping and compatibility with downstream tools that expect SDRF-formatted metadata.

Analysis Modes

The --mode parameter selects the analysis engine:

Mode Description Engine
diann DIA quantitative proteomics DIA-NN
fragpipe DDA LFQ, TMT, or generic FragPipe workflows FragPipe

DIA Mode (--mode diann)

DIA workflows use DIA-NN for library-free or library-based DIA analysis. DIA-specific variants (e.g., phospho) are applied via -c conf/variants/*.config files (see DIA Method Variants).

nextflow run nf-core/msproteomics \
  --mode diann \
  --input samplesheet.csv \
  --database /path/to/database.fasta \
  --outdir results \
  -profile docker

If --database is omitted, the UniProt reference proteome for --organism (default Homo sapiens) is downloaded at runtime from conf/reference_proteomes.config; see Database Preparation. --database /path/to/database.fasta always takes precedence.

FragPipe Mode (--mode fragpipe)

FragPipe mode supports DDA LFQ, TMT label check, TMT quantification, and generic FragPipe workflows. The specific workflow is controlled by the --fragpipe_workflow and --tmt_mode parameters.

DDA LFQ -- uses MSFragger, MSBooster, Percolator, ProteinProphet, Philosopher, and IonQuant:

nextflow run nf-core/msproteomics \
  --mode fragpipe \
  --fragpipe_container your-registry/fragpipe:24.0 \
  --fragpipe_workflow LFQ-MBR.workflow \
  --input samplesheet.csv \
  --database /path/to/uniprot_human.fasta \
  --outdir results \
  -profile docker

TMT Label Check -- assess TMT labeling efficiency:

nextflow run nf-core/msproteomics \
  --mode fragpipe \
  --fragpipe_container your-registry/fragpipe:24.0 \
  --tmt_mode labelcheck \
  --input samplesheet.csv \
  --database /path/to/uniprot_human.fasta \
  --outdir results \
  -profile docker

TMT Quantification -- full TMT isobaric quantification:

nextflow run nf-core/msproteomics \
  --mode fragpipe \
  --fragpipe_container your-registry/fragpipe:24.0 \
  --tmt_mode quant \
  --input samplesheet.csv \
  --database /path/to/uniprot_human.fasta \
  --outdir results \
  -profile docker

Generic FragPipe -- run any FragPipe workflow by providing a .workflow configuration file:

nextflow run nf-core/msproteomics \
  --mode fragpipe \
  --fragpipe_container your-registry/fragpipe:24.0 \
  --fragpipe_workflow /path/to/custom.workflow \
  --input samplesheet.csv \
  --database /path/to/uniprot_human.fasta \
  --outdir results \
  -profile docker

Understanding .workflow Files

FragPipe .workflow files are key=value configuration files that control which tools run and their parameters. The pipeline ships three built-in workflow files in assets/:

File Use Case Key Tools
LFQ-MBR.workflow DDA label-free quantification with match-between-runs MSFragger, MSBooster, Percolator, ProteinProphet, IonQuant
TMT-labelcheck-229.workflow TMT label check for TMT6/10/11 (mass 229 Da) MSFragger, Percolator, ProteinProphet
TMT-labelcheck-304.workflow TMT label check for TMTpro/16/18 (mass 304 Da) MSFragger, Percolator, ProteinProphet

To create a custom .workflow file:

  1. Open the FragPipe GUI and configure your analysis
  2. Save the workflow configuration (File > Export Workflow)
  3. Provide the exported file to the pipeline with --fragpipe_workflow /path/to/custom.workflow

Key parameters users commonly tune in .workflow files:

  • Enzyme: trypsin, lysc, chymotrypsin, etc.
  • Missed cleavages: typically 1-2
  • Variable modifications: oxidation (M), acetylation (protein N-term)
  • Mass tolerances: precursor and fragment mass tolerances
  • Calibrate mass: 0 (disabled), 1 (single pass), 2 (two pass, recommended)

Note

TMT label check workflows use TMT as a variable modification (not fixed) to detect unlabeled peptides. This is intentional and differs from standard TMT quantification workflows.

The --fragpipe_mode parameter controls execution mode:

  • pipeline (default): Modular Nextflow subworkflow execution
  • headless: Single-process FragPipe headless execution

Key Parameters

Parameter Description Default
--input Path to CSV samplesheet Required
--mode Analysis mode: diann or fragpipe Required
--outdir Output directory Required
--database FASTA protein database UniProt reference proteome for --organism
--fragpipe_workflow FragPipe .workflow file or workflow name (FragPipe mode) None
--tmt_mode TMT analysis mode: labelcheck or quant (FragPipe mode) None
--fragpipe_mode FragPipe execution mode: pipeline or headless pipeline

Container Requirements

DIA Workflows

DIA workflows use publicly available containers for DIA-NN and associated tools. No special licensing is required.

FragPipe-Based Workflows (DDA LFQ, TMT, Generic)

The pipeline defaults to the public base image docker.io/fcyucn/fragpipe:24.0, which includes open-source FragPipe tools (Philosopher, Percolator, MSBooster, PTMShepherd, etc.).

However, commercial/licensed tools (MSFragger, IonQuant, DiaTracer) are not included in the public image. You must build your own container image with these tools under your own license terms. See Docker build instructions for academic and commercial container builds.

FragPipe is free for academic use. Commercial users must obtain a license from Fragmatics.

Profiles

Use -profile to select a software packaging method:

Profile Description
docker Run with Docker containers (recommended)
singularity Run with Singularity containers
apptainer Run with Apptainer containers
podman Run with Podman containers
conda Run with Conda environments (last resort)
test Minimal stub test for CI validation (DIA mode)
test_dia DIA workflow test with small DIA dataset
test_dda_lfq DDA LFQ workflow test with small DDA dataset
test_tmt TMT label check workflow test
test_tmtq TMT quantification workflow test

Multiple profiles can be combined: -profile test_dia,docker

Running the Pipeline

Basic Execution

nextflow run nf-core/msproteomics \
  --mode diann \
  --input samplesheet.csv \
  --database /path/to/database.fasta \
  --outdir results \
  -profile docker

Using a Parameters File

nextflow run nf-core/msproteomics -profile docker -params-file params.yaml
input: "samplesheet.csv"
mode: "fragpipe"
fragpipe_container: "your-registry/fragpipe:24.0"
fragpipe_workflow: "LFQ-MBR.workflow"
database: "/path/to/uniprot_human.fasta"
outdir: "results"

Resuming a Failed Run

nextflow run nf-core/msproteomics \
  --mode diann \
  --input samplesheet.csv \
  --database /path/to/database.fasta \
  --outdir results \
  -profile docker \
  -resume

Updating the Pipeline

nextflow pull nf-core/msproteomics

Reproducibility

Specify a pipeline version with -r to ensure reproducible results:

nextflow run nf-core/msproteomics -r 1.0.0 \
  --mode diann \
  --input samplesheet.csv \
  --database /path/to/database.fasta \
  --outdir results \
  -profile docker

Configuration Hierarchy

The pipeline loads configuration files automatically based on the --mode parameter:

DIA workflows (--mode diann):

  1. nextflow.config -- root config, always loaded
  2. conf/base_configs/diann.config -- DIA-NN computed params and helper functions
  3. conf/variants/*.config -- variant-specific overrides (e.g., diaphos) applied via -c
  4. conf/instruments/*.config -- instrument-specific overrides applied via -c

FragPipe workflows (--mode fragpipe):

  1. nextflow.config -- root config, always loaded (all params defined here)
  2. --tmt_type selects the TMT plex (TMT6, TMT10, TMT11, TMT16, TMT18, TMTPRO); the correct labelcheck workflow file is auto-selected
  3. conf/instruments/*.config -- instrument-specific overrides applied via -c

Custom Configuration

Resource Requests

To change compute resources for specific processes, see the nf-core resource tuning documentation.

Custom Tool Arguments

To pass additional arguments to specific tools, use the ext.args directive in a custom config file. See the nf-core tool arguments documentation.

Resource Recommendations

Resource requirements depend on dataset size and analysis type:

Dataset Size Files Recommended Memory Notes
Small < 20 16 GB Default resources sufficient
Medium 20-100 32 GB Increase MSFragger and IonQuant memory
Large 100+ 64 GB Consider splitting into batches
  • DIA-NN library generation is the most memory-intensive step; allocate 32-64 GB for large proteomes or phosphoproteomics
  • MSFragger memory scales with database size; large databases (>100,000 sequences) may need 32+ GB
  • FragPipe headless mode runs all tools in a single process and requires memory sufficient for the largest tool
  • FragPipe pipeline mode distributes work across processes, allowing smaller per-process memory allocation

For cloud environments, memory-optimized instances are recommended (e.g., AWS r6i family, GCP n2-highmem).

Warning

Do not use -c <file> to specify pipeline parameters. Custom config files specified with -c must only be used for resource tuning, output directories, or module arguments (ext.args).

Institutional Configs

The pipeline dynamically loads configurations from nf-core/configs. Check if your institution has a pre-configured profile available.

Running on Seqera Platform (Tower)

Nextflow Version

Set the Nextflow version in your launch pre-run script:

export NXF_VER=25.10.2

Compute Environments

FragPipe processes require compute environments with sufficient memory. Recommended instance families: r6i (memory-optimized).

Instrument-Specific Settings

DIA-NN in-silico library generation uses different m/z ranges and mass accuracy settings depending on the mass spectrometer. The pipeline ships per-instrument config files in conf/instruments/ that override the defaults when included with -c.

Abbreviation Instrument Config File diann_min_pr_mz diann_max_pr_mz diann_min_fr_mz diann_max_fr_mz diann_library_mass_acc diann_library_ms1_acc
ASC Thermo Ascend conf/instruments/thermo_ascend.config (default) 350 1050 200 1800 18 5
FLX Bruker timsTOF Flex conf/instruments/bruker_flex.config 100 1700 100 1700 15 15

Usage example:

nextflow run nf-core/msproteomics \
  --mode diann \
  --input samplesheet.csv \
  --database /path/to/database.fasta \
  --outdir results \
  -c conf/instruments/bruker_flex.config \
  -profile docker

Thermo Ascend users can omit the -c flag since the defaults are already optimized for Ascend instruments.

DIA Method Variants

Specialized DIA analysis modes are available as variant config files in conf/variants/. These configs adjust DIA-NN parameters (e.g., variable modifications, phosphosite monitoring) for specific experimental designs.

Variant Config File Description
DIA Phosphoproteomics conf/variants/diaphos.config Phosphosite localization with UniMod:21 monitoring

Usage example:

nextflow run nf-core/msproteomics \
  --mode diann \
  --input samplesheet.csv \
  --database /path/to/database.fasta \
  --outdir results \
  -c conf/variants/diaphos.config \
  -profile docker

Variant configs can be combined with instrument configs by specifying multiple -c flags:

nextflow run nf-core/msproteomics \
  --mode diann \
  --input samplesheet.csv \
  --database /path/to/database.fasta \
  --outdir results \
  -c conf/instruments/bruker_flex.config \
  -c conf/variants/diaphos.config \
  -profile docker

Nextflow Memory Requirements

To limit Nextflow's JVM memory usage, add the following to your environment:

NXF_OPTS='-Xms1g -Xmx4g'

Further Reading