metabolomics-workbench-database
Query Metabolomics Workbench REST API (4,200+ NIH studies) for metabolite ID, study discovery, RefMet standardization, m/z precursor searches, and gene/protein annotations.
Install / Use
npx skills add jaechang-hits/SciAgent-Skills --skill metabolomics-workbench-databaseInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Data & AnalyticsSupported Platforms
Our assessment of metabolomics-workbench-database
metabolomics-workbench-database scores 91/100 on our quality scale, 205th of 573 Data & Analytics skills we index (top 36%).
Its SKILL.md is 21 KB long, well organised into 47 sections with 17 code examples: a thorough specification that gives an agent plenty to work with.
It has 367 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 37 days ago, so metabolomics-workbench-database is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
metabolomics-workbench-database compared with similar skills
All 4 of these similar skills score higher than metabolomics-workbench-database; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| metabolomics-workbench-database (this skill)by jaechang-hits | 91 | 367 | 37d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 90.8k | 19d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.4k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.7k | today | MCP Server |
| crawl4aiby unclecode | 100 | 84.8k | 9d ago | MCP Server |
Frequently asked questions
- How do I install metabolomics-workbench-database?
- Run
npx skills add jaechang-hits/SciAgent-Skills --skill metabolomics-workbench-database. The install tabs above show the steps for each supported agent. - Which AI agents does metabolomics-workbench-database work with?
- It is written for Cursor, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is metabolomics-workbench-database safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is metabolomics-workbench-database still maintained?
- The repository was last updated 37 days ago, so metabolomics-workbench-database is actively maintained.
Skill content
View source on GitHubname: "metabolomics-workbench-database"
description: "Query Metabolomics Workbench REST API (4,200+ NIH studies) for metabolite ID, study discovery, RefMet standardization, m/z precursor searches, and gene/protein annotations. Quirks: compound input_item rejects name (use pubchem_cid/kegg_id/inchi_key/etc.); free-text → compound is a two-step refmet/match→refmet/name flow; moverz endpoint returns TSV text, not JSON. Use hmdb-database for local XML; pubchem-compound-search for general compound lookup."
license: "CC-BY-4.0"
Metabolomics Workbench Database — REST API Access
Overview
The Metabolomics Workbench (MW) REST API at https://www.metabolomicsworkbench.org/rest/ exposes 4,200+ metabolomics studies hosted at UCSD under NIH Common Fund sponsorship. URL pattern is /{context}/{input_item}/{input_value}/{output_item}/{format}. Contexts include compound, refmet, moverz, study, analysis, metabolite, gene, protein. Notable quirks discovered live:
compound/name/{x}is rejected —nameis not an allowed input_item. Usepubchem_cid,kegg_id,inchi_key,hmdb_id,regno,lm_id,formula,smiles, orabbrev. For free-text input, go throughrefmet/match/{x}first.refmet/name/{x}/allrequires the exact RefMet name (e.g.Glucose, notD-glucose); userefmet/match/{x}for fuzzy normalisation first.moverz/{REFMET|LIPIDS|MB}/{mz}/{ion}/{tol}/txtreturns TSV text (no JSON variant).- The
metstat/filter/...endpoint shown in older examples returns[]— replace withstudy/{context}/{value}/summary(or/metabolites) + client-side filtering.
No authentication required.
When to Use
- Searching metabolite records by PubChem CID, KEGG ID, InChIKey, HMDB ID, formula, or SMILES
- Discovering studies by species, disease, last_name, institute, analysis_type, or polarity
- Standardising metabolite names to RefMet nomenclature for cross-study integration
- Identifying unknown compounds from MS m/z values with adduct-aware matching (
moverz) - Retrieving experimental metabolite tables (analyses, abundances) from published studies
- Querying gene/protein annotations linked to metabolomics pathways
- Downloading raw mwTab files for local analysis
- For local 220K-metabolite XML parsing with NMR/MS spectra use
hmdb-databaseinstead - For live 110M-compound property lookups use
pubchem-compound-searchinstead
Prerequisites
- Python packages:
requests,pandas - No API key required: publicly accessible
- Rate limits: MW does not enforce strict limits; add
time.sleep(0.3)between bulk requests - Base URL:
https://www.metabolomicsworkbench.org/rest
pip install requests pandas
Quick Start
import requests
BASE = "https://www.metabolomicsworkbench.org/rest"
# Two-step free-text → compound (the API rejects compound/name/...)
def lookup_by_name(name):
# 1) Normalise to RefMet name
r = requests.get(f"{BASE}/refmet/match/{name}", timeout=30)
r.raise_for_status()
refmet = r.json()
if not refmet.get("refmet_name"):
return None
# 2) Pull full compound record by RefMet name (or by pubchem_cid)
r2 = requests.get(f"{BASE}/refmet/name/{refmet['refmet_name']}/all", timeout=30)
rec = r2.json() if r2.json() else {}
return rec if isinstance(rec, dict) else None
c = lookup_by_name("glucose")
print(f"{c['name']}: formula={c['formula']}, PubChem CID={c['pubchem_cid']}, "
f"InChIKey={c['inchi_key']}")
# Glucose: formula=C6H12O6, PubChem CID=5793, InChIKey=WQZGKKKJIJFFOK-GASJEMHNSA-N
Core API
Module 1: Compound Queries
compound/{input_item}/{input_value}/all/json — input_item must be one of regno, formula, inchi_key, lm_id, pubchem_cid, hmdb_id, kegg_id, smiles, abbrev. The legacy name input is rejected by the server.
import requests
BASE = "https://www.metabolomicsworkbench.org/rest"
# By PubChem CID
r = requests.get(f"{BASE}/compound/pubchem_cid/5793/all/json", timeout=30)
glucose = r.json()
print(f"PubChem 5793 -> {glucose['name']}, formula={glucose['formula']}, "
f"HMDB={glucose.get('hmdb_id')}, KEGG={glucose.get('kegg_id')}")
# By KEGG ID
r = requests.get(f"{BASE}/compound/kegg_id/C00031/all/json", timeout=30)
print("KEGG C00031 ->", r.json()["name"])
# By InChIKey
r = requests.get(f"{BASE}/compound/inchi_key/WQZGKKKJIJFFOK-GASJEMHNSA-N/all/json", timeout=30)
print("InChIKey -> regno:", r.json()["regno"])
# Compound by formula returns a paged dict (multiple matches)
import requests
BASE = "https://www.metabolomicsworkbench.org/rest"
r = requests.get(f"{BASE}/compound/formula/C6H12O6/all/json", timeout=30)
matches = r.json()
print(f"Compounds with formula C6H12O6: {len(matches)}")
for k in list(matches)[:3]:
print(f" regno={matches[k]['regno']} name={matches[k]['name']}")
Module 2: Study Discovery
study/{input_item}/{input_value}/{output} — input_item includes study_id, study_title, last_name, institute, analysis_id, metabolite_id, kegg_id, refmet_name. output includes summary, metabolites, factors, data, available_studies, species, disease. summary for study_id returns a dict (keyed by accession when multiple); for last_name/institute it returns a list.
import requests, pandas as pd
BASE = "https://www.metabolomicsworkbench.org/rest"
# Single-study summary — `study/study_id/{id}/summary` returns a flat dict
# (keys: study_id, study_title, species, institute, analysis_type, ...)
r = requests.get(f"{BASE}/study/study_id/ST000001/summary", timeout=30)
s = r.json()
print(f"{s['study_id']}: {s['study_title'][:60]}")
print(f" Species : {s.get('species')} Institute: {s.get('institute')}")
print(f" Submit : {s.get('submission_date')}")
# Studies that detected a metabolite — `study/refmet_name/{x}/summary` returns
# a thin index of (refmet_name, kegg_id, study_id) rows. Chain study_id → full summary
# to get title and species.
import requests, pandas as pd
BASE = "https://www.metabolomicsworkbench.org/rest"
r = requests.get(f"{BASE}/study/refmet_name/Glucose/summary", timeout=60)
d = r.json()
rows = list(d.values()) if isinstance(d, dict) else d
print(f"Studies referencing 'Glucose': {len(rows)}")
print(pd.DataFrame(rows).head(5).to_string(index=False))
# refmet_name kegg_id study_id
# Glucose C00031 ST000001
# Glucose C00031 ST000002
# ...
Module 3: RefMet Standardisation
refmet/match/{user_text} is a fuzzy normaliser — returns the standard RefMet record (no pubchem_cid/kegg_id though). refmet/name/{exact_refmet_name}/all returns the full record including IDs. Use them as a two-step pipeline.
import requests
BASE = "https://www.metabolomicsworkbench.org/rest"
def normalise_to_refmet(user_text):
r = requests.get(f"{BASE}/refmet/match/{user_text}", timeout=30)
r.raise_for_status()
m = r.json()
if not m or not m.get("refmet_name"):
return None
return m["refmet_name"]
def refmet_full(refmet_name):
r = requests.get(f"{BASE}/refmet/name/{refmet_name}/all", timeout=30)
r.raise_for_status()
rec = r.json()
return rec if isinstance(rec, dict) and rec else None
name = normalise_to_refmet("alpha-D-glucose") # -> 'Glucose'
print(f"Normalised: {name}")
rec = refmet_full(name)
print(f" PubChem CID : {rec['pubchem_cid']}")
print(f" InChIKey : {rec['inchi_key']}")
print(f" Super class : {rec['super_class']} / {rec['main_class']} / {rec['sub_class']}")
Module 4: Study Filtering (replaces broken metstat)
The older metstat/filter/... endpoint returns []. Use the study context endpoints with client-side filtering instead.
import requests, pandas as pd
BASE = "https://www.metabolomicsworkbench.org/rest"
def study_ids_for_metabolite(refmet_name):
"""Return the study_id list that report a given RefMet name."""
r = requests.get(f"{BASE}/study/refmet_name/{refmet_name}/summary", timeout=60)
r.raise_for_status()
d = r.json()
rows = list(d.values()) if isinstance(d, dict) else d
return sorted({row["study_id"] for row in rows if row.get("study_id")})
def study_summary(study_id):
"""Pull full summary (title, species, institute, dates) for one study_id.
Response is a flat dict with keys study_id/study_title/species/institute/..."""
return requests.get(f"{BASE}/study/study_id/{study_id}/summary", timeout=30).json()
# Find glucose studies, then enrich the first few
ids = study_ids_for_metabolite("Glucose")
print(f"Studies referencing 'Glucose': {len(ids)}")
rows = [study_summary(sid) for sid in ids[:5]]
df = pd.DataFrame(rows)
print(df[["study_id", "study_title", "species"]].head().to_string(index=False))
Module 5: m/z Precursor Search (moverz)
moverz/{REFMET|LIPIDS|MB}/{mz}/{ion}/{tolerance}/txt returns tab-separated text (not JSON). The first DB selector (REFMET, LIPIDS, or MB) is required — mz as the first segment is rejected.
import requests, io, pandas as pd
BASE = "https://www.metabolomicsworkbench.org/rest"
def moverz_search(db, mz, ion, tolerance=0.005):
"""Search precursor m/z in REFMET / LIPIDS / MB and return a DataFrame.
Response is TSV text — no JSON variant."""
assert db in {"REFMET", "LIPIDS", "MB"}
r = requests.get(f"{BASE}/moverz/{db}/{mz}/{ion}/{tolerance}/txt", timeout=30)
r.raise_for_status()
return pd.read_csv(io.StringIO(r.text), sep="\t")
df = moverz_search("REFMET", 180.063, "M+H", 0.005)
print(f"Candidates at m/z 180.063 [M+H]+ (REFMET): {len(df)}")
print(df.head(5).to_string(index=False))
# Same query against the LIPIDS database
import requests, io, pandas as pd
BASE = "https://www.metabolomicsworkbench.org/rest"
r = requests.get(f"{BASE}/moverz/LIPIDS/760.585/M+H/0.01/txt", timeout=30)
df_lipids = pd.read_csv(io.StringIO(r.text), sep="\t")
print(f"Lipid candidates at m/z 760.585: {len(df_lipids)}")
print(df_lipids.head(3).to_string(index=False))
Module 6: Genes and Proteins
import requests
BASE = "https://www.metabolomicsworkbench.org/rest"
r = requests.get(f"{BASE}/gene/gene_symbol/HMGCR/all", timeout=30)
g = r.json()
print(f"{g['gene_symbol']} -> MGP: {g.get('mgp_id')} ({g.get('gene_name', '')[:60]})")
# Protein by UniProt accession
r2 = requests.get(f"{BASE}/protein/uniprot_id/P04035/all", timeout=30)
p = r2.json()
print(f"UniProt P04035: {p.get('protein_name', '')[:60]} organism={p.get('organism')}")
Key Concepts
Allowed input_item Values per Context
| Context | Valid input_item | Notes |
|---------|------------------|-------|
| compound | regno, formula, inchi_key, lm_id, pubchem_cid, hmdb_id, kegg_id, smiles, abbrev | name is rejected — go via refmet/match first |
| refmet | match, name, formula, exactmass, inchi_key, pubchem_cid, regno | match is fuzzy; name requires the canonical RefMet name |
| study | study_id, study_title, last_name, institute, analysis_id, metabolite_id, kegg_id, refmet_name | summary for study_id is a dict keyed by accession; for last_name/institute it's a list |
| moverz | (path) REFMET / LIPIDS / MB | First segment is the DB, not mz |
| gene | gene_id, gene_symbol, gene_name, mgp_id | Returns a dict |
| protein | mgp_id, gene_id, uniprot_id, gene_symbol | Returns a dict |
Output Type Conventions
output=summaryreturns a dict when the input identifier is unique (e.g.study_id), a list when it isn't (e.g.last_name).- Appending
/jsontostudy/.../summaryflips the response to TSV. Omit the format suffix — JSON is the default. moverzonly emits/txt(TSV); there is no JSON variant.
refmet/match vs refmet/name
refmet/match/{user_text}— fuzzy. Always returns a single dict withrefmet_name,formula, `ex
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
90.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.4kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Scrapling
85.7k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
crawl4ai
84.8kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
