pubmed-database
Programmatic PubMed access via NCBI E-utilities REST API. Covers Boolean/MeSH queries, field-tagged search, endpoints (ESearch, EFetch, ESummary, EPost, ELink), history server for batches, citation matching, systematic review strategies. Use for biomedical literature search or automated pipelines.
Install / Use
npx skills add jaechang-hits/SciAgent-Skills --skill pubmed-databaseInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Our assessment of pubmed-database
pubmed-database scores 91/100 on our quality scale, 1112th of 2,866 Automation skills we index (top 39%).
Its SKILL.md is 17 KB long, well organised into 64 sections with 13 code examples: a thorough specification that gives an agent plenty to work with.
It has 367 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 37 days ago, so pubmed-database is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
pubmed-database compared with similar skills
All 4 of these similar skills score higher than pubmed-database; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| pubmed-database (this skill)by jaechang-hits | 91 | 367 | 37d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 90.8k | 19d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.4k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.7k | today | MCP Server |
| crawl4aiby unclecode | 100 | 84.8k | 9d ago | MCP Server |
Frequently asked questions
- How do I install pubmed-database?
- Run
npx skills add jaechang-hits/SciAgent-Skills --skill pubmed-database. The install tabs above show the steps for each supported agent. - Which AI agents does pubmed-database work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is pubmed-database safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is pubmed-database still maintained?
- The repository was last updated 37 days ago, so pubmed-database is actively maintained.
Skill content
View source on GitHubname: pubmed-database description: >- Programmatic PubMed access via NCBI E-utilities REST API. Covers Boolean/MeSH queries, field-tagged search, endpoints (ESearch, EFetch, ESummary, EPost, ELink), history server for batches, citation matching, systematic review strategies. Use for biomedical literature search or automated pipelines. license: CC-BY-4.0
PubMed Database
Overview
PubMed is the U.S. National Library of Medicine's database providing free access to 36M+ biomedical citations from MEDLINE and life sciences journals. This skill covers programmatic access via the E-utilities REST API and advanced search query construction using Boolean operators, MeSH terms, and field tags.
When to Use
- Searching biomedical literature with structured Boolean/MeSH queries
- Building automated literature monitoring or extraction pipelines
- Conducting systematic literature reviews or meta-analyses
- Retrieving article metadata, abstracts, or citation information by PMID/DOI
- Finding related articles or exploring citation networks programmatically
- Batch processing large sets of PubMed records
- For Python-native PubMed access, prefer BioPython (
Bio.Entrez) — this skill covers direct REST API usage - For broader academic search (non-biomedical), use OpenAlex or Semantic Scholar APIs
Prerequisites
pip install requests # HTTP client for E-utilities API
# Optional: pip install biopython — Bio.Entrez wrapper (higher-level API)
API Rate Limits:
- Without API key: 3 requests/second
- With API key: 10 requests/second (register at https://www.ncbi.nlm.nih.gov/account/)
- Always include
User-Agentheader with contact email
Quick Start
import requests
import time
BASE_URL = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils/"
API_KEY = "YOUR_API_KEY" # Optional but recommended
def pubmed_request(endpoint, params):
"""Reusable helper for E-utilities API calls with rate limiting."""
params.setdefault("api_key", API_KEY)
response = requests.get(f"{BASE_URL}{endpoint}", params=params)
response.raise_for_status()
time.sleep(0.1 if API_KEY != "YOUR_API_KEY" else 0.34) # Rate limit
return response
# Search → Fetch workflow
search_resp = pubmed_request("esearch.fcgi", {
"db": "pubmed", "term": "CRISPR[tiab] AND 2024[dp]",
"retmax": 5, "retmode": "json"
})
pmids = search_resp.json()["esearchresult"]["idlist"]
print(f"Found {len(pmids)} articles: {pmids}")
fetch_resp = pubmed_request("efetch.fcgi", {
"db": "pubmed", "id": ",".join(pmids),
"rettype": "abstract", "retmode": "text"
})
print(fetch_resp.text[:500])
Core API
1. Search Query Construction
Build PubMed queries using Boolean operators, field tags, and MeSH terms.
# Boolean operators: AND, OR, NOT (must be uppercase)
queries = {
"basic": "diabetes AND treatment AND 2024[dp]",
"synonyms": "(metformin OR insulin) AND type 2 diabetes",
"exclude": "cancer NOT review[pt]",
"phrase": '"gene expression" AND RNA-seq',
"field_tags": "smith ja[au] AND cancer[tiab] AND 2023[dp]",
}
# Common field tags:
# [tiab] = title/abstract [au] = author [mh] = MeSH term
# [pt] = publication type [dp] = date [ta] = journal
# [1au] = first author [lastau] = last author
# [affil] = affiliation [doi] = DOI [pmid] = PubMed ID
# Date filtering
date_queries = {
"single_year": "cancer AND 2024[dp]",
"range": "cancer AND 2020:2024[dp]",
"specific": "cancer AND 2024/03/15[dp]",
}
# MeSH terms — controlled vocabulary for precise searching
mesh_queries = {
# [mh] includes narrower terms automatically
"broad": "diabetes mellitus[mh]",
# [majr] limits to major topic focus
"focused": "diabetes mellitus[majr]",
# MeSH + subheading
"therapy": "diabetes mellitus, type 2[mh]/drug therapy",
# Substance name
"drug": "metformin[nm] AND diabetes mellitus[mh]",
}
# Common MeSH subheadings:
# /diagnosis /drug therapy /epidemiology /etiology
# /prevention & control /therapy /genetics
2. ESearch — Search and Retrieve PMIDs
# Basic search
resp = pubmed_request("esearch.fcgi", {
"db": "pubmed",
"term": "CRISPR[tiab] AND genome editing[tiab] AND 2024[dp]",
"retmax": 100,
"retmode": "json",
"sort": "relevance", # or "pub_date", "first_author"
})
result = resp.json()["esearchresult"]
pmids = result["idlist"]
total = result["count"]
print(f"Total hits: {total}, Retrieved: {len(pmids)}")
# With history server (for large result sets > 500)
resp = pubmed_request("esearch.fcgi", {
"db": "pubmed",
"term": "cancer AND 2024[dp]",
"usehistory": "y",
"retmode": "json",
})
result = resp.json()["esearchresult"]
webenv = result["webenv"]
query_key = result["querykey"]
total = int(result["count"])
print(f"Stored {total} results on history server")
3. EFetch — Download Full Records
# Fetch abstracts as text
resp = pubmed_request("efetch.fcgi", {
"db": "pubmed",
"id": ",".join(pmids[:10]),
"rettype": "abstract",
"retmode": "text",
})
print(resp.text[:500])
# Fetch XML for structured parsing
resp = pubmed_request("efetch.fcgi", {
"db": "pubmed",
"id": ",".join(pmids[:10]),
"rettype": "xml",
"retmode": "xml",
})
# Fetch from history server (batch processing)
batch_size = 500
for start in range(0, total, batch_size):
resp = pubmed_request("efetch.fcgi", {
"db": "pubmed",
"query_key": query_key,
"WebEnv": webenv,
"retstart": start,
"retmax": batch_size,
"rettype": "xml",
"retmode": "xml",
})
print(f"Fetched records {start}–{start + batch_size}")
time.sleep(0.5) # Extra delay for large batches
4. ESummary and ELink — Summaries and Related Articles
# ESummary — lightweight document summaries
resp = pubmed_request("esummary.fcgi", {
"db": "pubmed",
"id": ",".join(pmids[:5]),
"retmode": "json",
})
for uid, data in resp.json()["result"].items():
if uid == "uids":
continue
print(f"PMID {uid}: {data.get('title', '')[:80]}")
print(f" Journal: {data.get('fulljournalname', '')}, "
f"Date: {data.get('pubdate', '')}")
# ELink — find related articles
resp = pubmed_request("elink.fcgi", {
"dbfrom": "pubmed",
"db": "pubmed",
"id": pmids[0],
"cmd": "neighbor",
"retmode": "json",
})
# Links to other NCBI databases
resp = pubmed_request("elink.fcgi", {
"dbfrom": "pubmed",
"db": "pmc", # PubMed Central
"id": pmids[0],
"retmode": "json",
})
5. Citation Matching and Identifier Lookup
# Search by identifiers
id_queries = {
"pmid": "12345678[pmid]",
"doi": "10.1056/NEJMoa123456[doi]",
"pmc": "PMC123456[pmc]",
}
# ECitMatch — match partial citations to PMIDs
# Format: journal|year|volume|first_page|author_name|key|
citation = "Science|2008|320|5880|1185|key1|"
resp = pubmed_request("ecitmatch.cgi", {
"db": "pubmed",
"rettype": "xml",
"bdata": citation,
})
print(f"Matched PMID: {resp.text.strip()}")
# Batch citation matching
citations = [
"Nature|2020|580|7801|71|ref1|",
"Science|2019|366|6463|347|ref2|",
]
resp = pubmed_request("ecitmatch.cgi", {
"db": "pubmed",
"rettype": "xml",
"bdata": "\r".join(citations),
})
6. Publication Filtering
# Filter by publication type
type_filters = {
"rcts": "randomized controlled trial[pt]",
"reviews": "systematic review[pt]",
"meta": "meta-analysis[pt]",
"guidelines": "guideline[pt]",
"case_reports": "case reports[pt]",
}
# Filter by text availability
availability = {
"free_text": "free full text[sb]",
"has_abstract": "hasabstract[text]",
}
# Combine filters
query = (
"diabetes mellitus[mh] AND "
"randomized controlled trial[pt] AND "
"2023:2024[dp] AND "
"free full text[sb] AND "
"english[la]"
)
resp = pubmed_request("esearch.fcgi", {
"db": "pubmed", "term": query, "retmax": 100, "retmode": "json"
})
print(f"Free RCTs on diabetes (2023-2024): {resp.json()['esearchresult']['count']}")
Key Concepts
E-utilities Endpoint Summary
| Endpoint | Purpose | Key Parameters |
|----------|---------|----------------|
| esearch.fcgi | Search, return PMIDs | term, retmax, sort, usehistory |
| efetch.fcgi | Download full records | id/query_key+WebEnv, rettype, retmode |
| esummary.fcgi | Lightweight summaries | id, retmode |
| epost.fcgi | Upload UIDs to server | id (comma-separated) |
| elink.fcgi | Related articles, cross-DB | id, dbfrom, db, cmd |
| einfo.fcgi | List databases/fields | db (optional) |
| egquery.fcgi | Count hits across DBs | term |
| espell.fcgi | Spelling suggestions | term |
| ecitmatch.cgi | Match citations to PMIDs | bdata |
History Server Pattern
For result sets >500 articles, use the history server to avoid URL length limits:
- ESearch with
usehistory=y→ returnsWebEnv+QueryKey - EFetch in batches using
WebEnv+QueryKey+retstart/retmax - EPost to upload additional PMIDs to the same
WebEnv
Automatic Term Mapping (ATM)
When no field tag is specified, PubMed maps terms through: MeSH Translation Table → Journals Translation Table → Author Index → Full Text. Bypass ATM with explicit field tags or quoted phrases.
Common MeSH Subheadings
| Subheading | Abbreviation | Use | |------------|-------------|-----| | /diagnosis | /DI | Diagnostic methods | | /drug therapy | /DT | Pharmaceutical treatment | | /epidemiology | /EP | Disease patterns | | /etiology | /ET | Disease causes | | /genetics | /GE | Genetic aspects | | /prevention & control | /PC | Preventive measures | | /therapy | /TH | Treatment approaches |
Common Workflows
Workflow 1: Systematic Review Search
import requests, time, json
# 1. Define PICO-structured query
query = (
# Population
"(diabetes mellitus, type 2[mh] OR type 2 diabetes[tiab]) AND "
# Intervention + Comparison
"(metformin[nm] OR lifestyle modification[tiab]) AND "
# Outcome
"(glycemic control[tiab] OR HbA1c[tiab]) AND "
# Study design filter
"(randomized controlled trial[pt] OR systematic review[pt]) AND "
# Date range
"2020:2024[dp]"
)
# 2. Search with history server
resp = pubmed_request("esearch.fcgi", {
"db": "pubmed", "term": query,
"usehistory": "y", "retmode": "json"
})
result = resp.json()["esearchresult"]
total = int(result["count"])
print(f"Systematic review hits: {total}")
# 3. Batch fetch all results as XML
import xml.etree.ElementTree as ET
articles = []
for start in range(0, total, 200):
resp = pubmed_request("efetch.fcgi", {
"db": "pubmed", "query_key": result["querykey"],
"WebEnv": result["webenv"],
"retstart": start, "retmax": 200,
"rettype": "xml", "retmode": "xml"
})
root = ET.fromstring(resp.text)
for article in root.findall('.//PubmedArticle'):
pmid = article.findtext('.//PMID')
title = article.findtext('.//ArticleTitle')
articles.append({"pmid": pmid, "title": title})
time.sleep(0.5)
print(f"Retrieved {len(articles)} articles for screening")
Workflow 2: Literature Monitoring Pipeline
import json, datetime
# 1. Construct monitoring query
topic_query = (
"(CRISPR[tiab] OR gene editing[tiab]) AND "
"(therapeutics[tiab] OR clinical trial[pt])"
)
# 2. Search recent publications (last 30 days)
today = datetime.date.today()
start_date = today - datetime.timedelta(days=30)
query = f"{topic_query} AND {start_date.strftime('%Y/%m/%d')}:{today.strftime('%Y/%m/%d')}[dp]"
resp = pubmed_request("esearch.fcgi", {
"db": "pubmed", "term": query,
"retmax": 100, "retmode": "json", "sort": "pub_date"
})
pmids = resp.json()["esearchresult"]["idlist"]
# 3. Get summaries for n
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
90.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.4kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Scrapling
85.7k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
crawl4ai
84.8kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
