pubchem-mcp-server
Search the PubChem chemical database for compounds, properties, safety data, bioactivity, cross-references, and entity summaries via MCP. STDIO or Streamable HTTP.
Install / Use
claude mcp add cyanheads -- npx -y github:cyanheads/pubchem-mcp-serverIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AI & Machine LearningSupported Platforms
Our assessment of pubchem-mcp-server
pubchem-mcp-server scores 84/100 on our quality scale, 654th of 960 AI & Machine Learning skills we index.
Its MCP Server is 19 KB long, well organised into 36 sections with 11 code examples: a thorough specification that gives an agent plenty to work with.
It has 10 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated 3 days ago, so pubchem-mcp-server is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 97/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
pubchem-mcp-server compared with similar skills
All 4 of these similar skills score higher than pubchem-mcp-server; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| pubchem-mcp-server (this skill)by cyanheads | 84 | 10 | 3d ago | MCP Server |
| claude-memby thedotmack | 100 | 99.1k | 1d ago | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 95.3k | 2d ago | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.8k | 1d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.9k | today | CLAUDE.md |
Frequently asked questions
- How do I install pubchem-mcp-server?
- Run
claude mcp add cyanheads -- npx -y github:cyanheads/pubchem-mcp-server. The install tabs above show the steps for each supported agent. - Which AI agents does pubchem-mcp-server work with?
- It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
- Is pubchem-mcp-server safe to use?
- It is Apache-2.0-licensed and scores 97/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is pubchem-mcp-server still maintained?
- The repository was last updated 3 days ago, so pubchem-mcp-server is actively maintained.
Skill content
View source on GitHubPublic Hosted Server: https://pubchem.caseyjhand.com/mcp
</div>Overview
Chemical compound and bioassay data from PubChem's PUG REST and PUG View APIs. Search compounds by identifier, formula, or structure; fetch physicochemical properties, safety data, bioactivity, interactions, cross-references, and 3D structures; find bioassays by biological target. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.
Tools
| Tool | Description |
|:---|:---|
| pubchem_search_compounds | Search for compounds by name, SMILES, InChIKey, formula, substructure, superstructure, or 2D similarity. |
| pubchem_get_compound_details | Get physicochemical properties, descriptions, synonyms, drug-likeness, and classification for compounds by CID. |
| pubchem_get_compound_image | Fetch a 2D structure diagram (PNG) for a compound by CID. |
| pubchem_get_compound_3d_structure | Fetch a 3D conformer (atomic coordinates and bonds) for a compound by CID, as parsed JSON or raw SDF. |
| pubchem_get_compound_xrefs | Get external database cross-references (PubMed, patents, genes, proteins, etc.). |
| pubchem_get_compound_safety | Get GHS hazard classification and safety data for one or more compounds by CID (batch). |
| pubchem_get_bioactivity | Get a compound's bioactivity profile: assay results, targets, and activity values; filter by outcome or molecular target. |
| pubchem_get_compound_interactions | Get drug-drug, drug-food, and chemical-target interactions for a compound by CID. |
| pubchem_search_assays | Find bioassays by biological target (gene symbol, protein, Gene ID, UniProt accession). |
| pubchem_get_summary | Get summaries for PubChem entities: assays, genes, proteins, taxonomy. |
Resources
Compound and assay records are also exposed as URI-templated resources, backed by the same client methods as the tools; many MCP clients are tool-only and never surface resources.
| Resource | Description |
|:---|:---|
| pubchem://compound/{cid} | Core physicochemical properties (JSON). |
| pubchem://compound/{cid}/safety | GHS hazard classification (JSON). |
| pubchem://compound/{cid}/image | 2D structure diagram (PNG). |
| pubchem://compound/{cid}/xrefs | External cross-references (JSON). |
| pubchem://compound/{cid}/bioactivity | Bioassay activity profile (JSON). |
| pubchem://assay/{aid} | BioAssay summary (JSON). |
Capability reference
pubchem_search_compounds <sub>tool</sub>
- Five search strategies: identifier (name/SMILES/InChIKey, batched 1-25), formula (Hill notation, optional
allowOtherElements), substructure/superstructure containment, or 2D Tanimoto similarity (threshold 70-100, default 90) - Each strategy needs its own fields — identifier:
identifierType+identifiers; formula:formula; substructure/superstructure/similarity:query+queryType— and a missing or blank one is rejected before the upstream call - Caps at 200 CIDs per page (default 20);
offsetpages to a ceiling of 10,000 — identifier lookups resolve every match up front so paging is free, while formula/structure/similarity searches cost more upstream per deep page - Optional
propertieshydration avoids a follow-uppubchem_get_compound_detailscall - Identifier mode reports
unresolvedIdentifiersfor inputs that resolved to no CID — no PubChem match, or a SMILES PubChem cannot interpret — while the rest of the batch still resolves, plus notices when multiple inputs collide on one CID - A query PubChem cannot search on (malformed SMILES or formula, a
*wildcard atom, a CID with no record) fails fast with asearch_query_rejectedhint naming what to fix - Reports an exact
totalFoundwhen the full match set was observed, or atotalFoundAtLeastfloor when a bounded upstream search saturated
pubchem_get_compound_details <sub>tool</sub>
- Up to 100 CIDs per call; 27 available properties, defaulting to a core set of 14 (formula, weight, IUPAC name, SMILES forms, InChIKey, XLogP, TPSA, H-bond/rotatable-bond counts, heavy atom count, charge, complexity)
- Optional textual descriptions, paged via
descriptionOffset/maxDescriptions(default 3, up to 20) — fetched only for the first 10 CIDs in the batch, remaining CIDs listed inskippedCids - Optional synonyms for every found CID, paged via
synonymOffset/maxSynonyms(default 20, up to 100) - Optional drug-likeness assessment (Lipinski Rule of Five + Veber rules), computed from the returned properties at no extra latency
- Optional pharmacological classification (FDA classes/mechanisms, MeSH classes, ATC codes) — same 10-CID fan-out cap as descriptions
- Per-CID
found: falsedistinguishes a nonexistent CID from a real compound PubChem simply has no data for
pubchem_get_compound_image <sub>tool</sub>
- Single CID;
sizeis"small"(100x100) or"large"(300x300, default) - Returns base64-encoded PNG plus width/height
- Typed
cid_not_founderror when PubChem has no record for the CID
pubchem_get_compound_3d_structure <sub>tool</sub>
- Single CID;
format="json"(default) returns parsed atoms (element + x/y/z) and bonds,format="sdf"returns the raw V2000 SDF text maxAtoms/maxBondscap the JSON preview (default 200 each);atomCount/bondCountalways report the full totals, with any capping disclosed via enrichmentincludeRawSdfbypasses the default 500-line cap on the raw SDF text- Optional
includeAlternateConformerIdslists conformer IDs beyond the default - Typed
no_3d_structureerror when PubChem has no computed 3D coordinates (large molecules, mixtures, some salts)
pubchem_get_compound_xrefs <sub>tool</sub>
- Single CID; one or more
xrefTypes— string IDs (RegistryID,RNfor CAS numbers,PatentID) and numeric IDs (PubMedID,GeneID,ProteinGI,TaxonomyID) - Paged per type:
maxPerTypeup to 500 (default 50), with the sameoffsetapplied across every requested type - Each type reports its own
totalAvailableandtruncatedflag - Empty-result notice distinguishes "this compound has none of the requested types" from a possibly-mistyped CID
pubchem_get_compound_safety <sub>tool</sub>
- Batch of 1-25 CIDs
- Returns GHS signal word, pictograms, hazard statements (H-codes), and precautionary statements (P-codes), with source attribution
- Per-CID
status:ok,no_ghs_data(compound exists, no deposited classification), orcid_not_found(no PubChem record at all) — kept distinct so a bad CID never reads as "no hazards on file" - Precautionary statements carry a
decodedflag — false for codes needing label-specific fill text or outside the decoder table; the code itself is still authoritative
pubchem_get_bioactivity <sub>tool</sub>
- Single CID; filter by
outcomeFilter(active/inactive/all, defaultall) and/ortargetGeneId/targetAccession - Caps at 100 results per page (default 20);
offsetreaches the rest - Reports
totalAssays/activeCount/inactiveCountfor the whole compound, plusfilteredCount/returnedCountfor the current page - Notices distinguish "no bioactivity data at all" from "the filter excluded everything" from "offset past the end"
pubchem_get_compound_interactions <sub>tool</sub>
- Single CID; one or more
kinds—drug-drug(DrugBank),drug-food,target(binding/activity from BindingDB, ChEMBL, and others); default["drug-drug"] maxEntriesper kind per page (1-50, default 10);offsetcounts source records rather than returned entries, capped at 2,147,483,646- Each kind pages independently —
paging[]reports per-kindtotalRecords/nextOffset/truncated; the top-levelnextOffsetis populated only when exactly one requested kind still has records left - A kind that fails to retrieve is named in
failedKindswithout failing the kinds that succeeded
pubchem_search_assays <sub>tool</sub>
- Search by
targetType:genesymbol/proteinname(text),geneid(NCBI Gene ID),proteinaccession(UniProt) - Caps at 200 AIDs per page (default 50);
offsetpages to the total found - Rejects a blank
targetQueryand a non-numericgeneidquery before the upstream call - Reports
totalFoundacross all pages and distinguishes "no match" from "offset past the end"
pubchem_get_summary <sub>tool</sub>
entityType:assay(AID),gene(NCBI Gene ID),protein(UniProt accession), ortaxonomy(Tax ID); up to 10 identifiers per call- Per-identifier
foundflag; populated fields depend onentityType(taxonomy includes an orderedlineage, gene includessymbol/taxonomy) - Notice reports how many identifiers were not found and which ID type
entityTypeexpects
pubchem://compound/{cid} <sub>resource</sub>
- Core physicochemical properties (the same default 14-property set as
pubchem_get_compound_details), asapplication/json - Throws a typed not-found when the CID doesn't exist in PubChem
- Use
pubchem_get_compound_detailsto select specific properties or add descriptions, synonyms, drug-likeness, and classification
pubchem://compound/{cid}/safety <sub>resource</sub>
- GHS hazard classification as
application/json status(ok/no_ghs_data/cid_not_found) is the only signal distinguishing a bad CID from a compound with no deposited classification — a resource read has no notice surface
pubchem://compound/{cid}/image <sub>resource</sub>
- 2D structure diagram, 300x300 PNG, returned as a base64 blob
- Use
pubchem_get_compound_imagefor the 100x100 size option
pubchem://compound/{cid}/xrefs <sub>resource</sub>
- Fo
Truncated for display — read the full file on GitHub.
Related Skills
claude-mem
99.1kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
95.3kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Understand-Anything
85.8kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.9kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
