SkillAgentSearch skills...

maxquant-proteomics

MaxQuant + Perseus proteomics pipeline: run MaxQuant for LFQ and SILAC; parse proteinGroups.txt in Python; filter contaminants/decoys; log2 + median-normalize; impute MNAR; t-test with FDR; volcano plot; GO/pathway enrichment.

Install / Use

npx skills add jaechang-hits/SciAgent-Skills --skill maxquant-proteomics

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

91/100

Category

Automation

Supported Platforms

Universal

Our assessment of maxquant-proteomics

maxquant-proteomics scores 91/100 on our quality scale, 1101st of 2,866 Automation skills we index (top 39%).

Its SKILL.md is 28 KB long, well organised into 57 sections with 18 code examples: a thorough specification that gives an agent plenty to work with.

It has 367 GitHub stars, a meaningful sign that others use it.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
11/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 37 days ago, so maxquant-proteomics is actively maintained.
  • No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
  • Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.

Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.

maxquant-proteomics compared with similar skills

All 4 of these similar skills score higher than maxquant-proteomics; compare them before choosing.

SkillScoreStarsUpdatedFormat
maxquant-proteomics (this skill)by jaechang-hits9136737d agoSKILL.md
Agent-Reachby Panniantong10090.8k19d agoCLAUDE.md
headroomby headroomlabs-ai10074.4ktodayCLAUDE.md
Scraplingby D4Vinci10085.7ktodayMCP Server
crawl4aiby unclecode10084.8k9d agoMCP Server

Frequently asked questions

How do I install maxquant-proteomics?
Run npx skills add jaechang-hits/SciAgent-Skills --skill maxquant-proteomics. The install tabs above show the steps for each supported agent.
Which AI agents does maxquant-proteomics work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is maxquant-proteomics safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is maxquant-proteomics still maintained?
The repository was last updated 37 days ago, so maxquant-proteomics is actively maintained.

name: "maxquant-proteomics" description: "MaxQuant + Perseus proteomics pipeline: run MaxQuant for LFQ and SILAC; parse proteinGroups.txt in Python; filter contaminants/decoys; log2 + median-normalize; impute MNAR; t-test with FDR; volcano plot; GO/pathway enrichment. Use Proteome Discoverer for Thermo-native processing; FragPipe/MSFragger for GPU-accelerated DB search." license: "Apache-2.0"

MaxQuant + Perseus — Proteomics Analysis Pipeline

Overview

MaxQuant is the community-standard software for label-free quantification (LFQ) and SILAC proteomics. It performs database search, protein grouping, and intensity-based quantification from raw LC-MS/MS files, producing proteinGroups.txt as the primary output. Downstream statistical analysis — filtering, normalization, imputation, differential abundance testing, and visualization — is performed in Python using pandas, scipy, and matplotlib/seaborn, mirroring the Perseus workflow in a reproducible scripting environment.

When to Use

  • Performing label-free quantification (LFQ) of proteins across multiple biological conditions — MaxQuant's MaxLFQ algorithm is the community benchmark
  • Running SILAC (stable isotope labeling) experiments with light/heavy or triple-label designs
  • Processing iTRAQ or TMT isobaric labeling experiments via MaxQuant's reporter ion quantification
  • Identifying and quantifying proteins when you need the widely-cited MaxQuant output format (proteinGroups.txt) for comparison with published datasets
  • Performing statistical differential abundance analysis on MaxQuant outputs without installing Perseus (GUI-only, Windows)
  • Generating publication-quality volcano plots and GO enrichment from proteomics data in a reproducible Python workflow
  • Use Proteome Discoverer instead when working with Thermo raw files requiring instrument-native processing or Sequest HT
  • Use FragPipe/MSFragger instead for GPU-accelerated database search (3–10× faster) or when processing DIA (data-independent acquisition) data
  • Use omics-plotting SKILL after differential abundance analysis or GSEA for publication-quality plots

Prerequisites

  • MaxQuant: Windows software; download from https://maxquant.org/ (v2.4+); requires .NET 6 runtime
  • Python packages: pandas, numpy, scipy, matplotlib, seaborn, statsmodels, gseapy
  • Data requirements: Thermo .raw files or mzML-converted files; FASTA protein database (UniProt reviewed + contaminant database)
  • Environment: MaxQuant runs on Windows (GUI or CLI); Python analysis runs cross-platform
pip install pandas numpy scipy matplotlib seaborn statsmodels gseapy
# Install pyMaxQuant for programmatic mqpar.xml configuration
pip install pymaxquant

Pre-flight Interview

Settle these with the user before writing any analysis code.

decisions:
  - id: D1
    param: quantificationStrategy
    kind: required
    source: user
    ask: "How were samples quantified - label-free, SILAC, or isobaric tags?"
    default: null

  - id: D2
    param: proteinSequenceDatabase
    kind: required
    source: user
    ask: "Which organism's protein FASTA, and which release, should spectra be searched against?"
    default: null

  - id: D3
    param: digestionEnzyme
    kind: required
    source: user
    ask: "Which protease was used, and how many missed cleavages should be allowed?"
    default: "trypsin, up to 2 missed cleavages"

  - id: D4
    param: modifications
    kind: required
    source: user
    ask: "Which modifications are fixed, and which variable - and is there an enrichment such as phospho to search for?"
    default: "fixed carbamidomethyl on cysteine; variable oxidation and N-terminal acetylation"

  - id: D5
    param: falseDiscoveryRates
    kind: required
    source: user
    ask: "What false-discovery rate should peptide and protein identifications be held to?"
    default: "1% at both levels"

  - id: D6
    param: matchBetweenRuns
    kind: required
    source: user
    depends_on: [D1]
    ask: "Should identifications be transferred between runs by retention time, recovering more proteins at the cost of some transferred errors?"
    default: "off"

  - id: D7
    param: lfqMinRatioCount
    kind: optional
    source: user
    depends_on: [D1]
    ask: "How many shared peptides must two samples have before their abundances are normalized against each other?"
    default: 2
    skip_if: "labelled quantification"

  - id: D8
    param: missingValueImputation
    kind: required
    source: user
    ask: "How should proteins missing from a sample be treated - left missing, or imputed as below the detection limit?"
    default: "imputed from a downshifted distribution"

  - id: D9
    param: differentialThresholds
    kind: required
    source: user
    ask: "What significance and fold-change cuts define a changed protein?"
    default: "FDR 0.05, log2 fold change 1.0"

D8 is where proteomics differs from RNA-seq and where results most often turn on an unexamined default. A protein absent from every control and present in every treated sample is the strongest possible result or an artefact of detection limits, and the imputation choice decides which one the volcano plot shows. D6 interacts with it: transferred identifications fill in exactly the values that would otherwise be imputed.

Quick Start

import pandas as pd
import numpy as np

# Load MaxQuant output
df = pd.read_csv("combined/txt/proteinGroups.txt", sep="\t", low_memory=False)
print(f"Raw protein groups: {len(df)}")

# Filter contaminants, reverse decoys, only-by-site
mask = (
    (df["Potential contaminant"] != "+") &
    (df["Reverse"] != "+") &
    (df["Only identified by site"] != "+")
)
df = df[mask].copy()
print(f"After filtering: {len(df)} protein groups")

# Extract LFQ intensity columns
lfq_cols = [c for c in df.columns if c.startswith("LFQ intensity ")]
print(f"LFQ columns: {lfq_cols}")

# Log2-transform (0 → NaN)
lfq = df[lfq_cols].replace(0, np.nan)
lfq = np.log2(lfq)
print(f"Valid values per sample:\n{lfq.notna().sum()}")

Workflow

Step 1: Configure MaxQuant Parameters via mqpar.xml

MaxQuant is controlled by an XML parameter file (mqpar.xml). Edit it programmatically to set file paths, enzyme, modifications, and quantification type before running the search.

import xml.etree.ElementTree as ET

def update_mqpar(template_path: str, output_path: str,
                 raw_files: list[str], fasta_path: str,
                 experiment_names: list[str]) -> None:
    """Update mqpar.xml with sample-specific file paths."""
    tree = ET.parse(template_path)
    root = tree.getroot()

    # Set raw file paths
    file_paths_node = root.find(".//filePaths")
    file_paths_node.clear()
    for rf in raw_files:
        elem = ET.SubElement(file_paths_node, "string")
        elem.text = rf

    # Set experiment names (maps files to conditions)
    experiments_node = root.find(".//experiments")
    experiments_node.clear()
    for name in experiment_names:
        elem = ET.SubElement(experiments_node, "string")
        elem.text = name

    # Set FASTA database
    fasta_node = root.find(".//fastaFiles/FastaFileInfo/fastaFilePath")
    fasta_node.text = fasta_path

    tree.write(output_path, xml_declaration=True, encoding="utf-8")
    print(f"Written: {output_path}")

# Example usage
raw_files = [
    r"C:\Data\ctrl_rep1.raw",
    r"C:\Data\ctrl_rep2.raw",
    r"C:\Data\treat_rep1.raw",
    r"C:\Data\treat_rep2.raw",
]
update_mqpar(
    template_path="mqpar_template.xml",
    output_path="mqpar.xml",
    raw_files=raw_files,
    fasta_path=r"C:\Databases\human_uniprot_contaminants.fasta",
    experiment_names=["ctrl", "ctrl", "treat", "treat"],
)

Key mqpar.xml parameters (set in template or edit directly):

<!-- Enzyme and search settings -->
<enzymes>
  <string>Trypsin/P</string>
</enzymes>
<maxMissedCleavages>2</maxMissedCleavages>
<variableModifications>
  <string>Oxidation (M)</string>
  <string>Acetyl (Protein N-term)</string>
</variableModifications>
<fixedModifications>
  <string>Carbamidomethyl (C)</string>
</fixedModifications>

<!-- LFQ settings -->
<lfqMode>1</lfqMode>                     <!-- 1 = LFQ enabled -->
<lfqMinRatioCount>2</lfqMinRatioCount>   <!-- minimum peptides for LFQ -->
<matchBetweenRuns>True</matchBetweenRuns>

<!-- FDR thresholds -->
<peptideFdr>0.01</peptideFdr>
<proteinFdr>0.01</proteinFdr>

Step 2: Run MaxQuant from Command Line (Windows)

MaxQuant can be run headlessly from the Windows command prompt using the bundled MaxQuantCmd.exe.

REM Windows Command Prompt — run MaxQuant with configured mqpar.xml
REM Adjust path to match your MaxQuant installation directory

set MQ_PATH=C:\Program Files\MaxQuant\bin\MaxQuantCmd.exe
set MQPAR=C:\Projects\proteomics\mqpar.xml

"%MQ_PATH%" "%MQPAR%"

REM For specific workflow steps only (useful for reruns):
REM Step IDs: 0=write tables, 1=feature detection, 7=peptide identification
"%MQ_PATH%" "%MQPAR%" --steps 1,7,11
# Cross-platform: run MaxQuant under Wine on Linux/macOS (CI/server use)
wine MaxQuantCmd.exe mqpar.xml

# Monitor progress log
tail -f combined/proc/#runningTimes.txt

Step 3: Load and Filter proteinGroups.txt

Filter out reverse decoys, potential contaminants, and proteins only identified by modification site.

import pandas as pd
import numpy as np

def load_protein_groups(path: str) -> pd.DataFrame:
    """Load MaxQuant proteinGroups.txt with quality filters applied."""
    df = pd.read_csv(path, sep="\t", low_memory=False)
    print(f"Total protein groups: {len(df)}")

    # Remove reverse decoys, contaminants, and only-by-site hits
    n_before = len(df)
    df = df[
        (df.get("Reverse", pd.Series("")) != "+") &
        (df.get("Potential contaminant", pd.Series("")) != "+") &
        (df.get("Only identified by site", pd.Series("")) != "+")
    ].copy()
    print(f"After quality filter: {len(df)} ({n_before - len(df)} removed)")

    # Parse gene names (take first entry for multi-gene groups)
    df["Gene names"] = df["Gene names"].fillna("Unknown").str.split(";").str[0]

    # Set unique index on majority protein ID
    df = df.set_index("Majority protein IDs")
    return df

# Load output
pg = load_protein_groups("combined/txt/proteinGroups.txt")

# Identify LFQ intensity columns
lfq_cols = [c for c in pg.columns if c.startswith("LFQ intensity ")]
print(f"LFQ samples ({len(lfq_cols)}): {lfq_cols}")
# Output: LFQ samples (6): ['LFQ intensity ctrl_1', 'LFQ intensity ctrl_2', ...]

Step 4: Log2 Transform and Median Normalize LFQ Intensities

Replace zero intensities with NaN (missing values in MaxQuant are exported as 0), log2-transform, then apply per-sample median centering.

def prepare_lfq_matrix(df: pd.DataFrame, lfq_cols: list[str]) -> pd.DataFrame:
    """Extract, transform, and normalize LFQ intensity matrix."""
    # Extract and rename columns (strip 'LFQ intensity ' prefix)
    lfq = df[lfq_cols].copy()
    lfq.columns = [c.replace("LFQ intensity ", "") for c in lfq_cols]

    # Replace 0 with NaN (MaxQuant encodes missing as 0)
    lfq = lfq.replace(0, np.nan)

    # Log2 transform
    lfq = np.log2(lfq)

    # Median centering per sample (subtract per-column median of valid values)
    col_medians = lfq.median(axis=0)
    global_median = col_medians.median()
    lfq = lfq.subtract(col_medians, axis=1).add(global_median)

    print(f"Matrix shape: {lfq.shape}")
    print(f"Missing values per sample:\n{lfq.isna().sum()}")
    print(f"Valid values per sample:\n{lfq.notna().sum()}")
    return lfq

lfq_matrix = prepare_lfq_matrix(pg, lfq_cols)
# Matrix shape: (3241, 6)
# Missing values per sample: ctrl_1: 421, ctrl_2: 389, ...

Step 5: Impute Missing Values (MNAR Strategy)

Missing-not-at-random (MNAR) values arise from proteins below the detection limit. Impute from the low end of the obser

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars367
CategoryAutomation
Updated1mo ago
Forks36

Languages

Python

Trust signals

88/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 medium