hmdb-database
Parse HMDB (Human Metabolome Database) local XML for metabolite info, chemical properties, biological context, disease links, spectra, and cross-DB mapping. No REST API — uses ~6 GB XML download. Use drugbank-database-access for drugs; pubchem-compound-search for live lookups.
Install / Use
npx skills add jaechang-hits/SciAgent-Skills --skill hmdb-databaseInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Data & AnalyticsSupported Platforms
Our assessment of hmdb-database
hmdb-database scores 91/100 on our quality scale, 203rd of 573 Data & Analytics skills we index (top 36%).
Its SKILL.md is 24 KB long, well organised into 37 sections with 21 code examples: a thorough specification that gives an agent plenty to work with.
It has 367 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 37 days ago, so hmdb-database is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
hmdb-database compared with similar skills
All 4 of these similar skills score higher than hmdb-database; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| hmdb-database (this skill)by jaechang-hits | 91 | 367 | 37d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 90.8k | 19d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.4k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.7k | today | MCP Server |
| crawl4aiby unclecode | 100 | 84.8k | 9d ago | MCP Server |
Frequently asked questions
- How do I install hmdb-database?
- Run
npx skills add jaechang-hits/SciAgent-Skills --skill hmdb-database. The install tabs above show the steps for each supported agent. - Which AI agents does hmdb-database work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is hmdb-database safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is hmdb-database still maintained?
- The repository was last updated 37 days ago, so hmdb-database is actively maintained.
Skill content
View source on GitHubname: "hmdb-database" description: "Parse HMDB (Human Metabolome Database) local XML for metabolite info, chemical properties, biological context, disease links, spectra, and cross-DB mapping. No REST API — uses ~6 GB XML download. Use drugbank-database-access for drugs; pubchem-compound-search for live lookups." license: "CC-BY-4.0"
HMDB Database — Local XML Access
Overview
Query the Human Metabolome Database (HMDB, 220,000+ metabolite entries) by parsing locally downloaded XML with Python's ElementTree. Covers metabolite lookup, chemical properties, biological context (pathways, enzymes, biofluids), disease/biomarker associations, spectral data for metabolite identification, and cross-database ID mapping to KEGG, PubChem, ChEBI, and DrugBank.
When to Use
- Looking up metabolite information (description, chemical class, cellular location) by HMDB ID or name
- Retrieving chemical properties (molecular weight, formula, SMILES, InChI, logP, PSA) for metabolomics analysis
- Finding pathway associations and enzyme links for a set of metabolites
- Identifying biofluid/tissue locations of metabolites (blood, urine, CSF, saliva)
- Querying disease associations and normal/abnormal concentration ranges for biomarker discovery
- Extracting NMR or MS spectral peak lists for metabolite identification
- Mapping HMDB IDs to KEGG, PubChem, ChEBI, DrugBank, or other databases
- For drug-specific data (interactions, targets, pharmacology) use
drugbank-database-accessinstead - For live compound property queries without downloading use
pubchem-compound-searchinstead
Prerequisites
- HMDB XML download: Register at https://hmdb.ca/downloads — download
hmdb_metabolites.xml.zip(~6 GB uncompressed) - Python packages:
lxml(faster XPath) or standardxml.etree.ElementTree,pandas - No public REST API: HMDB has no programmatic REST API. All access is via local XML parsing or the web interface
- R package (optional):
hmdbQueryon CRAN provides some query wrappers but is limited and outdated - Rate limits: N/A for local XML parsing. The web interface has no documented rate limits but is not intended for scraping
pip install lxml pandas
Quick Start
import xml.etree.ElementTree as ET
NS = {'hmdb': 'http://www.hmdb.ca'}
tree = ET.parse('hmdb_metabolites.xml') # 60-120s for full XML
root = tree.getroot()
# Build lookup index (HMDB ID + lowercase name -> element)
metabolite_index = {}
for met in root.findall('hmdb:metabolite', NS):
accession = met.find('hmdb:accession', NS)
name = met.find('hmdb:name', NS)
if accession is not None and accession.text:
metabolite_index[accession.text] = met
if name is not None and name.text:
metabolite_index[name.text.lower()] = met
def find_metabolite(query):
"""Find metabolite by HMDB ID or name (case-insensitive)."""
return metabolite_index.get(query) or metabolite_index.get(query.lower())
met = find_metabolite('HMDB0000122') # Glucose
name = met.find('hmdb:name', NS).text
formula = met.find('hmdb:chemical_formula', NS).text
print(f"{name}: {formula}")
# Glucose: C6H12O6
Core API
1. XML Setup and Metabolite Lookup
import xml.etree.ElementTree as ET
NS = {'hmdb': 'http://www.hmdb.ca'}
tree = ET.parse('hmdb_metabolites.xml')
root = tree.getroot()
print(f"Total metabolite entries: {len(root.findall('hmdb:metabolite', NS))}")
# Total metabolite entries: ~220000+
For memory-constrained environments, use iterparse:
met_names = {}
for event, elem in ET.iterparse('hmdb_metabolites.xml', events=('end',)):
if elem.tag == '{http://www.hmdb.ca}metabolite':
acc = elem.find('{http://www.hmdb.ca}accession')
name = elem.find('{http://www.hmdb.ca}name')
if acc is not None and name is not None and acc.text and name.text:
met_names[acc.text] = name.text
elem.clear()
print(f"Parsed {len(met_names)} metabolites via iterparse")
2. Chemical Properties
def get_chemical_properties(met_element):
"""Extract chemical properties from a metabolite entry."""
def txt(path):
el = met_element.find(path, NS)
return el.text if el is not None and el.text else None
# Fields: accession, name, chemical_formula, average_molecular_weight,
# monisotopic_molecular_weight, smiles, inchi, inchikey, state, iupac_name
return {tag: txt(f'hmdb:{tag}') for tag in [
'accession', 'name', 'chemical_formula', 'average_molecular_weight',
'monisotopic_molecular_weight', 'smiles', 'inchi', 'inchikey', 'state']}
props = get_chemical_properties(find_metabolite('HMDB0000158')) # L-Tyrosine
print(f"{props['name']}: MW={props['average_molecular_weight']}, SMILES={props['smiles']}")
# Extract taxonomy / chemical classification (ClassyFire ontology)
def get_classification(met_element):
tax = met_element.find('hmdb:taxonomy', NS)
if tax is None:
return {}
def txt(tag):
el = tax.find(f'hmdb:{tag}', NS)
return el.text if el is not None and el.text else None
return {'kingdom': txt('kingdom'), 'super_class': txt('super_class'),
'class': txt('class'), 'sub_class': txt('sub_class'),
'direct_parent': txt('direct_parent')}
print(get_classification(find_metabolite('HMDB0000122')))
# {'kingdom': 'Organic compounds', 'super_class': 'Organooxygen compounds', ...}
3. Biological Context
def get_pathways(met_element):
"""Extract metabolic pathway associations."""
pathways = []
for pw in met_element.findall('hmdb:pathways/hmdb:pathway', NS):
name = pw.find('hmdb:name', NS)
smpdb_id = pw.find('hmdb:smpdb_id', NS)
kegg_id = pw.find('hmdb:kegg_map_id', NS)
pathways.append({
'name': name.text if name is not None else None,
'smpdb_id': smpdb_id.text if smpdb_id is not None else None,
'kegg_map_id': kegg_id.text if kegg_id is not None else None,
})
return pathways
for pw in get_pathways(find_metabolite('HMDB0000122'))[:5]:
print(f" {pw['smpdb_id']}: {pw['name']}")
def get_biolocations(met_element):
"""Extract biofluid, tissue, and cellular locations."""
bp = 'hmdb:biological_properties/hmdb:'
def texts(path):
return [el.text for el in met_element.findall(path, NS) if el.text]
return {'biofluids': texts(f'{bp}biospecimen_locations/hmdb:biospecimen'),
'tissues': texts(f'{bp}tissue_locations/hmdb:tissue'),
'cellular': texts(f'{bp}cellular_locations/hmdb:cellular')}
locs = get_biolocations(find_metabolite('HMDB0000122'))
print(f"Glucose biofluids: {locs['biofluids']}")
def get_enzymes(met_element):
"""Extract associated enzymes/proteins with UniProt IDs."""
return [{'name': (p.find('hmdb:name', NS).text
if p.find('hmdb:name', NS) is not None else None),
'uniprot_id': (p.find('hmdb:uniprot_id', NS).text
if p.find('hmdb:uniprot_id', NS) is not None else None),
'gene_name': (p.find('hmdb:gene_name', NS).text
if p.find('hmdb:gene_name', NS) is not None else None)}
for p in met_element.findall('hmdb:protein_associations/hmdb:protein', NS)]
enz = get_enzymes(find_metabolite('HMDB0000122'))
print(f"Glucose-associated proteins: {len(enz)}")
for e in enz[:3]:
print(f" {e['gene_name']} ({e['uniprot_id']}): {e['name']}")
4. Disease and Biomarker Queries
def get_diseases(met_element):
"""Extract disease associations with OMIM IDs and PubMed references."""
diseases = []
for d in met_element.findall('hmdb:diseases/hmdb:disease', NS):
name = d.find('hmdb:name', NS)
omim = d.find('hmdb:omim_id', NS)
pmids = [r.find('hmdb:pubmed_id', NS).text
for r in d.findall('hmdb:references/hmdb:reference', NS)
if r.find('hmdb:pubmed_id', NS) is not None and r.find('hmdb:pubmed_id', NS).text]
diseases.append({'name': name.text if name is not None else None,
'omim_id': omim.text if omim is not None else None,
'pubmed_ids': pmids})
return diseases
diseases = get_diseases(find_metabolite('HMDB0000122'))
print(f"Glucose disease associations: {len(diseases)}")
for d in diseases[:3]:
print(f" {d['name']} (OMIM: {d['omim_id']})")
def get_concentrations(met_element, biospecimen='Blood'):
"""Extract normal/abnormal concentration data filtered by biospecimen."""
result = {'normal': [], 'abnormal': []}
for ctype, key in [('normal_concentrations', 'normal'),
('abnormal_concentrations', 'abnormal')]:
for c in met_element.findall(f'hmdb:{ctype}/hmdb:concentration', NS):
bio = c.find('hmdb:biospecimen', NS)
if bio is None or bio.text != biospecimen:
continue
def txt(tag):
el = c.find(f'hmdb:{tag}', NS)
return el.text if el is not None else None
result[key].append({'value': txt('concentration_value'),
'units': txt('concentration_units'),
'condition': txt('subject_condition')})
return result
conc = get_concentrations(find_metabolite('HMDB0000122'), 'Blood')
print(f"Glucose blood: {len(conc['normal'])} normal, {len(conc['abnormal'])} abnormal")
5. Spectral Data
def get_ms_spectra(met_element):
"""Extract MS/MS spectral peak lists (m/z + intensity)."""
spectra = []
for spec in met_element.findall('hmdb:spectra/hmdb:spectrum', NS):
stype = spec.find('hmdb:type', NS)
if stype is None or 'MS' not in (stype.text or ''):
continue
peaks = [{'mz': float(p.find('hmdb:mass_charge', NS).text),
'intensity': float(p.find('hmdb:intensity', NS).text or 0)}
for p in spec.findall('hmdb:ms_ms_peaks/hmdb:ms_ms_peak', NS)
if p.find('hmdb:mass_charge', NS) is not None
and p.find('hmdb:mass_charge', NS).text]
spectra.append({'type': stype.text, 'num_peaks': len(peaks), 'peaks': peaks})
return spectra
spectra = get_ms_spectra(find_metabolite('HMDB0000122'))
print(f"Glucose MS spectra: {len(spectra)}")
def get_nmr_spectra(met_element):
"""Extract NMR spectral peak lists (1H, 13C). Same pattern as MS above."""
spectra = []
for spec in met_element.findall('hmdb:spectra/hmdb:spectrum', NS):
stype = spec.find('hmdb:type', NS)
if stype is None or 'NMR' not in (stype.text or ''):
continue
nucleus = spec.find('hmdb:nucleus', NS)
shifts = [float(p.find('hmdb:chemical_shift', NS).text)
for p in spec.findall('hmdb:nmr_one_d_peaks/hmdb:nmr_one_d_peak', NS)
if p.find('hmdb:chemical_shift', NS) is not None
and p.find('hmdb:chemical_shift', NS).text]
spectra.append({'type': stype.text,
'nucleus': nucleus.text if nucleus is not None else None,
'num_peaks': len(shifts), 'chemical_shifts': shifts})
return spectra
nmr = get_nmr_spectra(find_metabolite('HMDB0000122'))
print(f"Glucose NMR spectra: {len(nmr)}")
6. Cross-Database Mapping
def get_external_ids(met_element):
"""Extract cross-database identifiers (KEGG, PubChem, ChEBI, DrugBank, CAS, etc.)."""
fields = {'kegg_id': 'KEGG', 'pubchem_compound_id': 'PubChem', 'chebi_id': 'ChEBI',
'drugbank_id': 'DrugBank', 'chemspider_id': 'ChemSpider',
'cas_registry_number': 'CAS', 'biocyc_id': 'BioCyc', 'pdb_id': 'PDB',
'foodb_id': 'FooDB', 'metlin_id': 'METLIN'}
ids = {}
for tag, label in fields.items():
el = met_element.find(
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
90.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.4kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Scrapling
85.7k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
crawl4ai
84.8kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
