pyopenms-mass-spectrometry
MS data processing with PyOpenMS for LC-MS/MS proteomics and metabolomics — mzML/mzXML I/O, signal processing (smoothing, peak picking, centroiding), feature detection/linking, peptide/protein ID with FDR, untargeted metabolomics. Use matchms for simple spectral matching.
Install / Use
npx skills add jaechang-hits/SciAgent-Skills --skill pyopenms-mass-spectrometryInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Customer SupportSupported Platforms
Our assessment of pyopenms-mass-spectrometry
pyopenms-mass-spectrometry scores 91/100 on our quality scale, 133rd of 320 Customer Support skills we index (top 42%).
Its SKILL.md is 23 KB long, well organised into 79 sections with 20 code examples: a thorough specification that gives an agent plenty to work with.
It has 367 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 37 days ago, so pyopenms-mass-spectrometry is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
pyopenms-mass-spectrometry compared with similar skills
All 4 of these similar skills score higher than pyopenms-mass-spectrometry; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| pyopenms-mass-spectrometry (this skill)by jaechang-hits | 91 | 367 | 37d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 90.8k | 19d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.4k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.7k | today | MCP Server |
| crawl4aiby unclecode | 100 | 84.8k | 9d ago | MCP Server |
Frequently asked questions
- How do I install pyopenms-mass-spectrometry?
- Run
npx skills add jaechang-hits/SciAgent-Skills --skill pyopenms-mass-spectrometry. The install tabs above show the steps for each supported agent. - Which AI agents does pyopenms-mass-spectrometry work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is pyopenms-mass-spectrometry safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is pyopenms-mass-spectrometry still maintained?
- The repository was last updated 37 days ago, so pyopenms-mass-spectrometry is actively maintained.
Skill content
View source on GitHubname: pyopenms-mass-spectrometry description: MS data processing with PyOpenMS for LC-MS/MS proteomics and metabolomics — mzML/mzXML I/O, signal processing (smoothing, peak picking, centroiding), feature detection/linking, peptide/protein ID with FDR, untargeted metabolomics. Use matchms for simple spectral matching. license: BSD-3-Clause
PyOpenMS — Mass Spectrometry Analysis
Overview
PyOpenMS provides Python bindings to the OpenMS C++ library for computational mass spectrometry. It supports proteomics and metabolomics data processing including file I/O for 10+ MS formats, signal processing, feature detection, peptide/protein identification, and quantitative analysis across samples.
When to Use
- Processing raw LC-MS/MS data (mzML, mzXML) for proteomics or metabolomics
- Detecting chromatographic features and linking them across multiple samples
- Identifying peptides and proteins from MS/MS search engine results with FDR control
- Running untargeted metabolomics workflows (peak picking → feature detection → alignment → annotation)
- Converting between mass spectrometry file formats (mzML, mzXML, featureXML, idXML)
- Smoothing, filtering, and centroiding raw spectral data
- For simple spectral library matching and metabolite identification, use matchms instead
- For protein sequence analysis (not mass spec), use biopython instead
Prerequisites
uv pip install pyopenms numpy pandas matplotlib
- Python 3.8+; NumPy for peak array operations
- Input data: mzML files (standard MS format), FASTA databases (for identification)
- All algorithms follow a consistent pattern:
algo = Algorithm(); params = algo.getParameters(); params.setValue(...); algo.setParameters(params)
Pre-flight Interview
Settle these with the user before writing any analysis code.
decisions:
- id: D1
param: instrumentResolution
kind: required
source: data
ask: "Is this high-resolution data, and are the spectra already centroided or still in profile mode?"
default: null
- id: D2
param: signalToNoiseThreshold
kind: required
source: user
depends_on: [D1]
ask: "How far above noise must a peak rise to be kept during centroiding?"
default: 1.0
skip_if: "spectra already centroided by the acquisition software"
- id: D3
param: massTolerance
kind: required
source: data
depends_on: [D1]
ask: "What mass accuracy should feature detection and linking assume?"
default: "10 ppm"
- id: D4
param: retentionTimeTolerance
kind: required
source: data
ask: "How far apart in retention time may the same feature appear across runs?"
default: "100 s"
- id: D5
param: chargeStateRange
kind: required
source: user
ask: "Which charge states should features be searched for?"
default: "1 to 3 - widen for intact protein or metabolite work"
- id: D6
param: intensityNormalization
kind: optional
source: user
ask: "Should spectra be normalized before comparison, and to total ion current or to the base peak?"
default: "not normalized"
- id: D7
param: smoothingFilter
kind: optional_conditional
source: data
ask: "Does the signal need smoothing before peak picking, and with which filter width?"
default: "no smoothing"
D1 governs almost everything after it: profile data that skips centroiding produces one feature per scan point, and high-resolution tolerances applied to low-resolution data link features that are not the same compound. Both yield a full feature table.
Quick Start
import pyopenms as ms
# Load mzML file
exp = ms.MSExperiment()
ms.MzMLFile().load("sample.mzML", exp)
print(f"Spectra: {exp.getNrSpectra()}, Chromatograms: {exp.getNrChromatograms()}")
# Examine first spectrum
spec = exp.getSpectrum(0)
mz, intensity = spec.get_peaks()
print(f"MS level: {spec.getMSLevel()}, RT: {spec.getRT():.2f}s, Peaks: {len(mz)}")
# Quick preprocessing: smooth + centroid
gauss = ms.GaussFilter()
p = gauss.getParameters(); p.setValue("gaussian_width", 0.1); gauss.setParameters(p)
gauss.filterExperiment(exp)
picker = ms.PeakPickerHiRes()
centroided = ms.MSExperiment()
picker.pickExperiment(exp, centroided)
print(f"Centroided spectra: {centroided.getNrSpectra()}")
Core API
Module 1: File I/O & Data Access
Read and write mass spectrometry data in multiple formats.
import pyopenms as ms
# Read mzML (standard MS format)
exp = ms.MSExperiment()
ms.MzMLFile().load("data.mzML", exp)
# Indexed access for large files (memory-efficient)
loader = ms.IndexedMzMLFileLoader()
indexed_file = ms.OnDiscMSExperiment()
loader.load("large_data.mzML", indexed_file)
spec = indexed_file.getSpectrum(0) # Load single spectrum on demand
print(f"Total spectra: {indexed_file.getNrSpectra()}")
# Read identification results (idXML)
protein_ids, peptide_ids = [], []
ms.IdXMLFile().load("results.idXML", protein_ids, peptide_ids)
print(f"Peptide IDs: {len(peptide_ids)}, Protein IDs: {len(protein_ids)}")
# Read feature map
fm = ms.FeatureMap()
ms.FeatureXMLFile().load("features.featureXML", fm)
print(f"Features: {fm.size()}")
# Write mzML with compression
exp_out = ms.MSExperiment()
# ... populate experiment ...
ms.MzMLFile().store("output.mzML", exp_out)
# Read FASTA database
entries = []
ms.FASTAFile().load("database.fasta", entries)
print(f"Proteins in DB: {len(entries)}")
for e in entries[:3]:
print(f" {e.identifier}: {e.sequence[:30]}...")
Supported formats: mzML, mzXML, mzData (spectra); featureXML, consensusXML (features); idXML, mzIdentML, pepXML (identifications); TraML (transitions); mzTab (results); FASTA (sequences)
Module 2: Signal Processing & Peak Picking
Preprocess raw spectral data for downstream analysis.
import pyopenms as ms
exp = ms.MSExperiment()
ms.MzMLFile().load("raw.mzML", exp)
# Gaussian smoothing
gauss = ms.GaussFilter()
p = gauss.getParameters()
p.setValue("gaussian_width", 0.15) # m/z width
gauss.setParameters(p)
gauss.filterExperiment(exp)
# Savitzky-Golay smoothing (alternative)
sg = ms.SavitzkyGolayFilter()
p = sg.getParameters()
p.setValue("frame_length", 15) # Must be odd
sg.setParameters(p)
# sg.filterExperiment(exp) # Use one smoother, not both
# Peak picking (centroiding) — required before feature detection
picker = ms.PeakPickerHiRes()
p = picker.getParameters()
p.setValue("signal_to_noise", 1.0)
picker.setParameters(p)
centroided = ms.MSExperiment()
picker.pickExperiment(exp, centroided)
print(f"Raw peaks in spec 0: {exp.getSpectrum(0).size()}")
print(f"Centroided peaks: {centroided.getSpectrum(0).size()}")
# Normalization
normalizer = ms.Normalizer()
p = normalizer.getParameters()
p.setValue("method", "to_one") # "to_one" or "to_TIC"
normalizer.setParameters(p)
normalizer.filterPeakMap(centroided)
# Peak filtering — remove low-intensity noise
mower = ms.ThresholdMower()
p = mower.getParameters()
p.setValue("threshold", 100.0) # Minimum intensity
mower.setParameters(p)
mower.filterPeakMap(centroided)
# Baseline reduction
morph = ms.MorphologicalFilter()
p = morph.getParameters()
p.setValue("struc_elem_length", 3.0) # m/z window
morph.setParameters(p)
morph.filterExperiment(exp)
Module 3: Feature Detection & Linking
Detect chromatographic features and link them across samples.
import pyopenms as ms
# Load centroided data
exp = ms.MSExperiment()
ms.MzMLFile().load("centroided.mzML", exp)
# Feature detection (proteomics — centroided data)
ff = ms.FeatureFinder()
features = ms.FeatureMap()
seeds = ms.FeatureMap()
params = ms.FeatureFinder().getParameters("centroided")
ff.run("centroided", exp, features, params, seeds)
print(f"Detected {features.size()} features")
# Access feature properties
for f in features[:5]:
print(f" RT: {f.getRT():.1f}s, m/z: {f.getMZ():.4f}, "
f"intensity: {f.getIntensity():.0f}, quality: {f.getOverallQuality():.3f}")
# Feature linking across samples — align retention times first
aligner = ms.MapAlignmentAlgorithmPoseClustering()
p = aligner.getParameters()
p.setValue("max_num_peaks_considered", 1000)
aligner.setParameters(p)
# Link features into consensus map
linker = ms.FeatureGroupingAlgorithmQT()
p = linker.getParameters()
p.setValue("distance_RT:max_difference", 60.0) # seconds
p.setValue("distance_MZ:max_difference", 10.0) # ppm
linker.setParameters(p)
consensus = ms.ConsensusMap()
linker.group([features_sample1, features_sample2, features_sample3], consensus)
print(f"Consensus features: {consensus.size()}")
# Export to pandas for downstream analysis
import pandas as pd
df = consensus.get_df()
print(f"Consensus table: {df.shape}")
Module 4: Peptide & Protein Identification
Process search engine results with FDR control and protein inference.
import pyopenms as ms
# Load search engine results
protein_ids, peptide_ids = [], []
ms.IdXMLFile().load("search_results.idXML", protein_ids, peptide_ids)
# Examine peptide hits
for pep_id in peptide_ids[:3]:
print(f"Spectrum: RT={pep_id.getRT():.1f}, MZ={pep_id.getMZ():.4f}")
for hit in pep_id.getHits():
seq = hit.getSequence()
print(f" {seq} score={hit.getScore():.4f} charge={hit.getCharge()}")
# FDR filtering (target-decoy approach)
fdr = ms.FalseDiscoveryRate()
fdr.apply(peptide_ids)
# Filter at 1% FDR
filtered = []
for pep_id in peptide_ids:
hits = [h for h in pep_id.getHits() if h.getScore() <= 0.01]
if hits:
pep_id.setHits(hits)
filtered.append(pep_id)
print(f"Peptide IDs at 1% FDR: {len(filtered)}")
# Protein inference
inference = ms.BasicProteinInferenceAlgorithm()
inference.run(peptide_ids, protein_ids)
for prot_id in protein_ids:
for hit in prot_id.getHits()[:5]:
print(f"Protein: {hit.getAccession()}, score: {hit.getScore():.4f}")
# Peptide sequence handling
seq = ms.AASequence.fromString("PEPTIDER")
print(f"Molecular weight: {seq.getMonoWeight():.4f}")
print(f"Formula: {seq.getFormula()}")
# Modified sequence
mod_seq = ms.AASequence.fromString("PEPTM(Oxidation)DER")
print(f"Modified weight: {mod_seq.getMonoWeight():.4f}")
# Enzymatic digestion
digestor = ms.ProteaseDigestion()
digestor.setEnzyme("Trypsin")
digest = []
digestor.digest(ms.AASequence.fromString("MKWVTFISLLLLFSSAYSRGVFRR"), digest)
print(f"Tryptic peptides: {len(digest)}")
Module 5: Metabolomics Pipeline
Complete untargeted metabolomics workflow from raw data to feature table.
import pyopenms as ms
# Step 1: Load and centroid raw data
exp = ms.MSExperiment()
ms.MzMLFile().load("metabolomics_sample.mzML", exp)
picker = ms.PeakPickerHiRes()
centroided = ms.MSExperiment()
picker.pickExperiment(exp, centroided)
# Step 2: Feature detection for metabolomics (small molecules)
ff = ms.FeatureFinder()
features = ms.FeatureMap()
seeds = ms.FeatureMap()
params = ms.FeatureFinder().getParameters("centroided")
params.setValue("isotopic_pattern:charge_low", 1)
params.setValue("isotopic_pattern:charge_high", 3)
ff.run("centroided", centroided, features, params, seeds)
print(f"Detected features: {features.size()}")
# Step 3: Adduct detection (group related adducts)
decharger = ms.MetaboliteAdductDecharger()
p = decharger.getParameters()
p.setValue("potential_adducts", "H:+:0.6;Na:+:0.3;K:+:0.1") # Positive mode
decharger.setParameters(p)
# decharger.compute(features, feature_map_out, consensus_map_out)
# Step 4: RT alignment across samples
import pyopenms as ms
import pandas as pd
# Assuming feature maps from multiple samples
sample_files = ["sample1.featureXML", "sample2.featureXML", "sample3.featureXML"]
feature_maps = []
for f in sample_files:
fm = ms.FeatureMap()
ms.FeatureXMLFile().load(f, fm)
feature_maps.append(fm)
# Align retention times
aligner = ms.MapAlignmentAlgorithmPoseClustering()
aligner.setReference(0) # Use first sample as reference
# Step 5
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
90.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.4kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Scrapling
85.7k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
crawl4ai
84.8kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
