matchms-spectral-matching
MS spectral matching and metabolite ID with matchms. Import spectra (mzML, MGF, MSP, JSON), filter/normalize peaks, score similarity (cosine, modified cosine, fingerprint), build reproducible pipelines, identify unknowns vs spectral libraries. Use pyopenms for full LC-MS/MS proteomics.
Install / Use
npx skills add jaechang-hits/SciAgent-Skills --skill matchms-spectral-matchingInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Our assessment of matchms-spectral-matching
matchms-spectral-matching scores 91/100 on our quality scale, 1100th of 2,866 Automation skills we index (top 39%).
Its SKILL.md is 25 KB long, well organised into 72 sections with 15 code examples: a thorough specification that gives an agent plenty to work with.
It has 367 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 37 days ago, so matchms-spectral-matching is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
matchms-spectral-matching compared with similar skills
All 4 of these similar skills score higher than matchms-spectral-matching; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| matchms-spectral-matching (this skill)by jaechang-hits | 91 | 367 | 37d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 90.8k | 19d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.4k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.7k | today | MCP Server |
| crawl4aiby unclecode | 100 | 84.8k | 9d ago | MCP Server |
Frequently asked questions
- How do I install matchms-spectral-matching?
- Run
npx skills add jaechang-hits/SciAgent-Skills --skill matchms-spectral-matching. The install tabs above show the steps for each supported agent. - Which AI agents does matchms-spectral-matching work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is matchms-spectral-matching safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is matchms-spectral-matching still maintained?
- The repository was last updated 37 days ago, so matchms-spectral-matching is actively maintained.
Skill content
View source on GitHubname: matchms-spectral-matching description: MS spectral matching and metabolite ID with matchms. Import spectra (mzML, MGF, MSP, JSON), filter/normalize peaks, score similarity (cosine, modified cosine, fingerprint), build reproducible pipelines, identify unknowns vs spectral libraries. Use pyopenms for full LC-MS/MS proteomics. license: Apache-2.0
Matchms — Spectral Matching & Metabolite Identification
Overview
Matchms is a Python library for mass spectrometry data processing focused on spectral similarity calculation and compound identification. It provides multi-format I/O, 50+ spectrum filters for metadata harmonization and peak processing, 8 similarity scoring functions, and a pipeline framework for reproducible analytical workflows.
When to Use
- Identifying unknown metabolites by matching MS/MS spectra against reference libraries
- Computing spectral similarity scores (cosine, modified cosine, fingerprint-based)
- Processing and standardizing mass spectral data from multiple formats (mzML, MGF, MSP, JSON)
- Building reproducible spectral processing pipelines for quality control
- Harmonizing metadata across spectral databases (compound names, SMILES, InChI, adducts)
- Large-scale spectral library comparisons and duplicate detection
- For full LC-MS/MS proteomics workflows (feature detection, protein ID), use pyopenms instead
- For chemical structure similarity without mass spectra, use rdkit fingerprint comparison
Prerequisites
uv pip install matchms numpy pandas
# For chemical structure processing (SMILES, InChI, fingerprints):
uv pip install matchms[chemistry]
- Python 3.8+; NumPy for peak array operations
- Input: spectral data in MGF, MSP, mzML, mzXML, JSON, or pickle format
- Reference library in any supported format for matching
Pre-flight Interview
Settle these with the user before writing any analysis code.
decisions:
- id: D1
param: referenceLibrary
kind: required
source: user
ask: "Which spectral library should queries be matched against?"
default: null
- id: D2
param: similarityMeasure
kind: required
source: user
ask: "Should matching require the same precursor mass, or allow a mass shift so modified analogues still match?"
default: "cosine on matched peaks, precursor mass fixed"
- id: D3
param: peakMatchTolerance
kind: required
source: data
ask: "How close must two peaks be in m/z to count as the same fragment?"
default: "0.1 Da - tighten toward 0.005 for high-resolution data"
- id: D4
param: spectrumPreprocessing
kind: required
source: user
ask: "Which peaks should be discarded before scoring - low-intensity noise, peaks near the precursor, or spectra with too few peaks to be informative?"
default: "drop peaks under 1% relative intensity, require at least 10 peaks"
- id: D5
param: scoreCutoff
kind: required
source: user
ask: "What similarity score, and how many matched peaks, should a hit need before it is reported as an identification?"
default: null
- id: D6
param: peakWeighting
kind: optional
source: user
ask: "Should the score weight fragment m/z as well as intensity, favouring informative high-mass fragments?"
default: "intensity only"
- id: D7
param: fingerprintSettings
kind: optional_conditional
source: user
ask: "Which molecular fingerprint should structural similarity use, when comparing hits to known structures?"
default: "not computed"
D3 and D5 have no safe shared default because they trade off directly: a loose tolerance with a low score cutoff returns an identification for every query spectrum, all of them plausible-looking. Set the tolerance from the instrument's actual accuracy and the cutoff from what the library supports.
Quick Start
from matchms.importing import load_from_mgf
from matchms.filtering import default_filters, normalize_intensities
from matchms.filtering import select_by_relative_intensity, require_minimum_number_of_peaks
from matchms import calculate_scores
from matchms.similarity import CosineGreedy
# Load and process query spectra
queries = list(load_from_mgf("queries.mgf"))
queries = [default_filters(s) for s in queries]
queries = [normalize_intensities(s) for s in queries if s is not None]
queries = [require_minimum_number_of_peaks(s, n_required=5) for s in queries if s is not None]
# Load reference library
refs = list(load_from_mgf("library.mgf"))
refs = [default_filters(s) for s in refs]
refs = [normalize_intensities(s) for s in refs if s is not None]
# Calculate similarity scores
scores = calculate_scores(references=refs, queries=queries,
similarity_function=CosineGreedy(tolerance=0.1))
# Get best matches for first query
best = scores.scores_by_query(queries[0], sort=True)[:5]
for match, score_tuple in best:
print(f"Score: {score_tuple['score']:.3f}, Matches: {score_tuple['matches']}")
Core API
Module 1: Spectrum I/O
Import spectra from multiple file formats and export processed data.
from matchms.importing import (load_from_mgf, load_from_mzml, load_from_msp,
load_from_json, load_from_mzxml, load_from_pickle,
load_from_usi)
from matchms.exporting import save_as_mgf, save_as_msp, save_as_json, save_as_pickle
# Import from various formats (returns generators)
spectra_mgf = list(load_from_mgf("library.mgf"))
spectra_mzml = list(load_from_mzml("data.mzML"))
spectra_msp = list(load_from_msp("nist_library.msp"))
spectra_json = list(load_from_json("gnps_spectra.json"))
print(f"Loaded: MGF={len(spectra_mgf)}, mzML={len(spectra_mzml)}")
# Export processed spectra
save_as_mgf(spectra_mgf, "processed.mgf")
save_as_json(spectra_mgf, "processed.json")
save_as_pickle(spectra_mgf, "spectra.pickle") # Fast for intermediate results
# Pickle for large datasets (fastest I/O)
from matchms.importing import load_from_pickle
spectra = list(load_from_pickle("spectra.pickle"))
from matchms import Spectrum
import numpy as np
# Create spectrum manually
mz = np.array([100.0, 150.0, 200.0, 250.0, 300.0])
intensities = np.array([0.1, 0.5, 0.9, 0.3, 0.7])
metadata = {
"precursor_mz": 325.5,
"ionmode": "positive",
"compound_name": "Caffeine",
"smiles": "CN1C=NC2=C1C(=O)N(C(=O)N2C)C"
}
spectrum = Spectrum(mz=mz, intensities=intensities, metadata=metadata)
# Access spectrum data
print(f"Peaks: {spectrum.peaks.mz}")
print(f"Precursor: {spectrum.get('precursor_mz')}")
print(f"Name: {spectrum.get('compound_name')}")
# Visualize
spectrum.plot()
Module 2: Spectrum Filtering & Processing
Apply metadata harmonization and peak processing filters. Matchms provides 50+ filters.
from matchms.filtering import (
default_filters, normalize_intensities,
select_by_relative_intensity, select_by_mz,
require_minimum_number_of_peaks, reduce_to_number_of_peaks,
remove_peaks_around_precursor_mz, add_losses
)
# default_filters applies: metadata cleanup, charge correction, adduct parsing
spectrum = default_filters(spectrum)
# Peak normalization (max intensity → 1.0)
spectrum = normalize_intensities(spectrum)
# Filter peaks by relative intensity (remove noise below 1%)
spectrum = select_by_relative_intensity(spectrum, intensity_from=0.01, intensity_to=1.0)
# Filter peaks by m/z range
spectrum = select_by_mz(spectrum, mz_from=50.0, mz_to=500.0)
# Keep top N peaks only
spectrum = reduce_to_number_of_peaks(spectrum, n_max=50)
# Remove peaks near precursor (common contaminants)
spectrum = remove_peaks_around_precursor_mz(spectrum, mz_tolerance=17.0)
# Require minimum peaks for matching
spectrum = require_minimum_number_of_peaks(spectrum, n_required=5)
# Add neutral losses (useful for NeutralLossesCosine)
spectrum = add_losses(spectrum)
if spectrum is not None:
print(f"After filtering: {len(spectrum.peaks.mz)} peaks")
# Chemical annotation filters (require matchms[chemistry])
from matchms.filtering import (
derive_inchi_from_smiles, derive_inchikey_from_inchi,
derive_smiles_from_inchi, add_fingerprint,
repair_inchi_inchikey_smiles, require_valid_annotation
)
# Derive chemical identifiers from SMILES
spectrum = derive_inchi_from_smiles(spectrum)
spectrum = derive_inchikey_from_inchi(spectrum)
# Add molecular fingerprint for structural similarity
spectrum = add_fingerprint(spectrum, fingerprint_type="morgan", nbits=2048)
# Validate annotations
spectrum = require_valid_annotation(spectrum)
if spectrum is not None:
print(f"InChIKey: {spectrum.get('inchikey')}")
Module 3: Similarity Scoring
Compare spectra using multiple similarity metrics.
from matchms import calculate_scores
from matchms.similarity import (
CosineGreedy, CosineHungarian, ModifiedCosine,
NeutralLossesCosine, FingerprintSimilarity,
MetadataMatch, PrecursorMzMatch
)
# CosineGreedy — fast peak matching (greedy algorithm)
scores = calculate_scores(references=library, queries=unknowns,
similarity_function=CosineGreedy(tolerance=0.1))
# ModifiedCosine — accounts for precursor mass differences (best for analog search)
scores = calculate_scores(references=library, queries=unknowns,
similarity_function=ModifiedCosine(tolerance=0.1))
# CosineHungarian — optimal peak matching (slower but more accurate)
scores = calculate_scores(references=library, queries=unknowns,
similarity_function=CosineHungarian(tolerance=0.1))
# NeutralLossesCosine — similarity based on neutral loss patterns
scores = calculate_scores(references=library, queries=unknowns,
similarity_function=NeutralLossesCosine(tolerance=0.1))
# Access results
for i, query in enumerate(unknowns[:3]):
best_matches = scores.scores_by_query(query, sort=True)[:3]
print(f"\nQuery {i}: precursor_mz={query.get('precursor_mz')}")
for ref, score_tuple in best_matches:
print(f" {ref.get('compound_name', 'Unknown')}: "
f"score={score_tuple['score']:.3f}, matches={score_tuple['matches']}")
# FingerprintSimilarity — structural similarity (requires fingerprints)
from matchms.similarity import FingerprintSimilarity
scores = calculate_scores(references=library, queries=unknowns,
similarity_function=FingerprintSimilarity(
similarity_measure="jaccard"))
# PrecursorMzMatch — fast mass-based pre-filtering
from matchms.similarity import PrecursorMzMatch
scores = calculate_scores(references=library, queries=unknowns,
similarity_function=PrecursorMzMatch(tolerance=0.1))
# Multi-metric scoring: combine peak + structural similarity
cosine_scores = calculate_scores(references=library, queries=unknowns,
similarity_function=CosineGreedy(tolerance=0.1))
fp_scores = calculate_scores(references=library, queries=unknowns,
similarity_function=FingerprintSimilarity())
Module 4: Processing Pipelines
Build reusable, reproducible multi-step processing workflows.
from matchms import SpectrumProcessor
from matchms.filtering import (
default_filters, normalize_intensities,
select_by_relative_intensity, require_minimum_number_of_peaks,
remove_peaks_around_precursor_mz, add_losses
)
# Define reusable pipeline
pipeline = SpectrumProcessor([
default_filters,
normalize_intensities,
lambda s: select_by_relative_intensity(s, intensity_from=0.01),
lambda s: remove_peaks_around_precursor_mz(s, mz_tolerance=17.0),
lambda s: require_minimum_number_of_peaks(s, n_required=5),
add_losses
])
# Apply to all spectra (filters returning None remove the spectrum)
processed = [pipeline(s) for s in raw_spectra]
processed = [s for s in processed if s is not Non
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
90.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.4kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Scrapling
85.7k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
crawl4ai
84.8kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
