mcp-okn
Next Generation MCP service to query the NSF Open Knowledge Network Knowledge Graphs
Install / Use
claude mcp add sbl-sdsc -- npx -y github:sbl-sdsc/mcp-oknIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
Development & EngineeringSupported Platforms
Tags
Skill content
View source on GitHubmcp-okn
An MCP server for querying the federated SPARQL endpoint
(https://apps.okn.us/federation/sparql) over the
Proto-OKN knowledge graphs.
It lets an LLM discover which knowledge graphs are relevant (from the
okn-registry descriptions), then run
SPARQL queries scoped to one or more named graphs of the form
https://purl.org/okn/frink/kg/{shortname}.
About Proto-OKN
Proto-OKN — the Prototype Open Knowledge Network — is a National Science Foundation initiative (with NASA, NIH, the National Institute of Justice, NOAA, and the U.S. Geological Survey) that funds research teams to build a publicly accessible, interconnected set of data repositories and knowledge graphs. The graphs span domains such as health, the environment, criminal justice, space exploration, and supply-chain security, and are served together over the OKN federated SPARQL endpoint that this server queries. The okn-registry catalogs the participating knowledge graphs.
Components of mcp-okn
This repository ships one MCP server, two Agent Skills, and relies on two external literature MCPs — each with a distinct role:
| Component | Role | |---|---| | <img src="docs/overview/mcp-okn.png" width="72" alt="mcp-okn"><br>mcp-okn | Federated OKN query service: discovers the relevant Proto-OKN graphs, writes SPARQL scoped to named graphs, aligns across KGs via crosswalks, and returns grounded rows with provenance and transcripts. | | <img src="docs/overview/okn-report-style.png" width="72" alt="okn-report-style"><br>okn-report-style | Report & reproducibility skill: turns an OKN analysis into a polished, reproducible deliverable (Markdown / HTML / Excel / figures / maps), tracking sources, versions, queries, and caveats. | | <img src="docs/overview/okn-bioanalysis.png" width="72" alt="okn-bioanalysis"><br>okn-bioanalysis | Biomedical workflow skill: cross-KG analysis of genes, diseases, chemicals, and drugs — enrichment, ortholog projection, mechanistic maps — ranking hypotheses by evidence and calling the literature MCPs to validate. | | <img src="docs/overview/pubmed-paperclip.png" width="72" alt="PubMed and Paperclip"><br>PubMed & Paperclip | Literature evidence validation: search PubMed and full-text collections, corroborate OKN-derived claims against sources, add citations, and flag conflicts and uncertainty. |
Complementary by design: the skills augment mcp-okn, which queries the OKN;
okn-bioanalysis can call PubMed & Paperclip for literature evidence validation, and
okn-report-style turns the results into reproducible deliverables.
Examples
Example prompts
Once the server is configured in your MCP client (see Connecting your client), just ask in natural language — the assistant picks the graphs, writes the SPARQL, and combines the results for you. Some prompts to try:
- "List all Proto-OKN knowledge graphs as a table of shortname and description." — Result
- "List all verified crosswalks, grouped by domain, with an example of what each answers." — Result
- "For each crosswalk, list the join key and the SPARQL skeleton." — Result
- "Give a high-level overview of the spoke-genelab knowledge graph — its main classes and relationships — and draw the schema diagram." — Result
- "Which genes does rdkg associate with autism spectrum disorder?" — Result
- "What is the maximum PFAS measurement in each county?" — Result
- "How do I join spoke-okn and prokn? Show the verified recipe and shared identifier." — Result
- "Which knowledge graphs supply GO, pathway, or trait annotations for a gene I can join on Entrez?" — uses
find_context_sourcesto list every supplier with its join key and size - "What version of prokn is loaded, and when was it last updated?" — reads the
okn-voidprovenance viaget_kg_version - "Create a chat transcript of this analysis." — create a transcript in a downloadable Markdown file
- "Create a chat transcript of this analysis in PDF format." — the server returns Markdown and the client converts the
.mdto a.pdffile (Claude Desktop / claude.ai)
Crosswalk queries & transcripts
A crosswalk is a verified way to join two (or three) Proto-OKN knowledge graphs on a shared identifier — for example linking a disease in one graph to the genes another graph associates with it via a common MONDO or DOID id. Because the graphs are built by different teams on different ontologies, the value of the federation is in these connections: a crosswalk is an integration opportunity where a question one graph can't answer alone becomes answerable by combining two. This section catalogs the verified crosswalks and shows the queries that exercise them.
A visual map of the whole network — all 162 crosswalks across 35 graphs, drawn as
direct KG-to-KG edges (edge width ∝ log of the verified join count). Each crosswalk is
its own edge, so multiple crosswalks between the same pair of graphs fan out as parallel
arcs. Identifier-bridged joins (e.g. DOID↔MONDO via ubergraph, HGNC→Entrez via
wikidata) are shown as direct edges with the bridge noted in the label and the line
styled (dashed for an ubergraph bridge, dotted for wikidata, solid for a direct join).
▶ Click the image to open the interactive, zoomable network.
Two resources, each backed by live federated SPARQL joins verified to actually answer the question (biomedical claims checked against PubMed / Paperclip; geospatial and industrial joins against their authoritative shared standard):
- Proto-OKN Crosswalk Inventory — a single-page map of the verified crosswalks: the joined KGs, shared key, row count, and a one-line note on what each answers. Start here to see which graphs connect and on what identifier.
- Cross-KG crosswalk catalog — 328 example questions worked end-to-end, each with a full transcript (the live SPARQL and its results), across 15 domains (Anatomy & Cell Type, Chemicals, Disease & Phenotype, Earth Observation, Environmental Toxicology, Function & Pathways, Genes, Geospatial, Hydrology, Industry & Supply Chain, Justice & Public Safety, Proteins, Publications, Social Determinants & Services, Taxonomy). Every crosswalk is now worked twice — the inventory carries 324 questions — two for every one of the 162 crosswalks — and the catalog has a transcript behind each, plus four questions on two extra stems (a second example on the spoke-genelab×spoke-okn Entrez axis, and the three-way gene dossier whose clique row was retired).
Every catalog row links to a standalone, replayable transcript — the prompt, the answer, and every verbatim SPARQL query with its result.
These transcripts are produced by create_chat_transcript and can be re-run
against the endpoint with scripts/replay_transcript.py.
The catalog was generated by driving the model with the
crosswalk generation prompt — list every
list_crosswalks recipe, write two research questions per crosswalk, run and
verify each as live SPARQL, and validate the findings against the literature.
Case studies
Fourteen end-to-end analyses that federate many Proto-OKN graphs into a single evidence-backed map — five of a disease's biology (genes, variants, pathways/gene sets, drugs, altered-activity signatures, and clinical/biomarker features), four of environmental exposure and justice (PFAS source attribution, the bisphenol chemical exposome, cumulative environmental-justice burden across U.S. counties, and flood-mobilised contamination routed downstream through the stream network), one of urban scaling (how disease, mortality and crime scale with settlement size), one of wildlife sentinel surveillance (whether Florida's wild-animal record and its contaminant record overlap at all), one of supply-chain fragility (physical manufacturing capacity, regulated industrial burden, software dependency risk and community vulnerability in one frame), and one of research-infrastructure criticality (which Earth-observation instruments the climate-modelling record actually leans on, and what modelling would stop being able to check if one went dark) — each finding tagged with its source(s) and evidence kind, then ranked by cross-source agreement. Every case study ships an interactive HTML report, a reproducibility record preserving every verbatim SPARQL query, and an Excel workbook.
That last question is answered twice — once by claude-opus-5, once by
gpt-5.6-sol, from the same prompt against the same two graphs — so the two
runs can be read side by side.
Prerequisites for re-running a case study:
- The two Skills —
okn-bioanalysis(analysis) andokn-report-style(report format). - The PubMed and Paperclip MCP connectors — for the literature-comparison step only.
| Case study | Model | Report | Literature Comparison | Data | Reproducibility | Folder |
|---|---|---|---|---|---|---|
| Type 2 diabetes — 16 KGs | claude-opus-4-8 | HTML | md | xlsx | md | files |
| Alzheimer's disease — 8 KGs | claude-opus-4-8 | HTML | md | xlsx | md | files |
| Multiple sclerosis — 14 KGs | claude-opus-4-8 | HTML | md | xlsx | md | files |
| Spaceflight-induced bone loss — 8 KGs | claude-opus-4-8 | HTML | md | xlsx | md | files |
| Spaceflight-associated neuro-ocular syndrome (SANS) — 6 KGs | claude-opus-4-8 | HTML | md | xlsx | md | files |
| PFAS source prioritization — 5 KGs | claude-opus-4-8 | [HTML](https://sbl-sdsc.github.io/mcp-okn/docs/examples/PFAS/PFAS_report.h
Truncated for display — read the full file on GitHub.
Related Skills
momen-cursurrules-prompt-file
40.6kCursor rules for building custom frontends with Momen.app as headless BaaS with GraphQL API, actionflows, AI agents, and Stripe integration.
semiotic-react-dataviz-cursorrules-prompt-file
40.6kCursor rules for Semiotic data visualization library with 30+ chart types, MCP server, and AI-assisted chart generation.
Agent-Reach
71.9kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
ruflo
67.9k🌊 The original agent meta-harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, RAG integration, and native Claude Code / Codex / Hermes and many more Integrated

