scientific-literature-search
Systematic strategies for searching scientific literature across PubMed, arXiv, Google Scholar, and AI-assisted tools. Covers PICO framework for clinical questions, three-tiered search (database-specific, AI-assisted, content extraction), PubMed field tags and MeSH, boolean query construction, and f…
Install / Use
npx skills add jaechang-hits/SciAgent-Skills --skill scientific-literature-searchInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Data & AnalyticsSupported Platforms
Our assessment of scientific-literature-search
scientific-literature-search scores 91/100 on our quality scale, 212th of 573 Data & Analytics skills we index (top 37%).
Its SKILL.md is 22 KB long, well organised into 36 sections with 9 code examples: a thorough specification that gives an agent plenty to work with.
It has 367 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 37 days ago, so scientific-literature-search is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
scientific-literature-search compared with similar skills
All 4 of these similar skills score higher than scientific-literature-search; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| scientific-literature-search (this skill)by jaechang-hits | 91 | 367 | 37d ago | SKILL.md |
| algorithmic-artby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 13d ago | SKILL.md |
| ui-ux-pro-maxby nextlevelbuilder | 100 | 130.2k | 13d ago | SKILL.md |
Frequently asked questions
- How do I install scientific-literature-search?
- Run
npx skills add jaechang-hits/SciAgent-Skills --skill scientific-literature-search. The install tabs above show the steps for each supported agent. - Which AI agents does scientific-literature-search work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is scientific-literature-search safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is scientific-literature-search still maintained?
- The repository was last updated 37 days ago, so scientific-literature-search is actively maintained.
Skill content
View source on GitHubname: scientific-literature-search description: "Systematic strategies for searching scientific literature across PubMed, arXiv, Google Scholar, and AI-assisted tools. Covers PICO framework for clinical questions, three-tiered search (database-specific, AI-assisted, content extraction), PubMed field tags and MeSH, boolean query construction, and full-text extraction. Use when planning a literature search or choosing a search tier." license: CC-BY-4.0
Scientific Literature Search
Overview
Scientific literature search is the foundation of evidence-based research. A well-executed search maximizes recall (finding all relevant papers) while maintaining precision (avoiding irrelevant results). This guide provides a systematic approach that combines database-specific query strategies, AI-assisted synthesis, and direct content extraction, organized into a three-tiered framework that scales from targeted lookups to comprehensive landscape reviews.
Key Concepts
The PICO Framework
For clinical and biomedical questions, structure queries using the PICO framework:
- P (Population): Who are you studying? (e.g., "Diabetes Mellitus"[MeSH])
- I (Intervention): What treatment or exposure? (e.g., "Metformin"[MeSH])
- C (Comparison): What is the alternative? (e.g., placebo, standard care)
- O (Outcome): What result are you measuring? (e.g., "Cardiovascular Diseases"[MeSH])
PICO queries can be combined with publication type filters to target specific evidence levels:
"Diabetes Mellitus"[MeSH] AND "Metformin"[MeSH] AND "Cardiovascular Diseases"[MeSH] AND ("clinical trial"[Publication Type] OR "meta-analysis"[Publication Type])
Three-Tiered Search Strategy
Literature search is most effective when approached in tiers of increasing breadth:
Tier 1 -- Database-Specific Searches (Most Reliable)
Query established academic databases (PubMed, arXiv, Google Scholar) for peer-reviewed, indexed content. This is the most reliable tier and should always be the starting point.
- PubMed (via Biopython
Bio.Entrez): Primary database for biomedical and life science literature. Supports MeSH controlled vocabulary and advanced field tags. - arXiv (via the
arxivpackage): Preprint server for physics, mathematics, computer science, and quantitative biology. Results appear faster than peer-reviewed journals. - Google Scholar (via the
scholarlypackage): Broadest coverage across all academic disciplines. Note: has aggressive rate limits on automated queries.
Best for: finding specific papers, systematic reviews, clinical evidence, preprints.
Tier 2 -- AI-Assisted Web Search (Comprehensive)
Use the Claude API with the web_search_20250305 server-side tool to synthesize broader context, identify research trends, and surface recent developments not yet indexed in databases. Also use general web search (e.g. via the duckduckgo-search package) for protocols, tutorials, and software documentation.
Best for: understanding the research landscape, complex multi-faceted questions, finding recent developments, identifying key researchers.
Avoid for: specific paper lookups (use Tier 1), citation counts (use Google Scholar), systematic reviews requiring reproducibility, searches where exact query terms must be documented.
Tier 3 -- Direct Content Extraction (Deep Dive)
Extract and analyze full-text content, PDFs, and supplementary materials from identified papers using trafilatura (HTML article extraction), pypdf (PDF text), and the Crossref API (DOI → supplementary file URLs).
Best for: detailed methodology extraction, data retrieval, protocol identification, supplementary data access.
PubMed Field Tags
PubMed supports field-specific searching to improve precision:
| Tag | Description | Example |
|-----|-------------|---------|
| [MeSH] | Medical Subject Heading (controlled vocabulary) | "Neoplasms"[MeSH] |
| [Title] | Title field only | "CRISPR"[Title] |
| [Title/Abstract] | Title or abstract | "gene therapy"[Title/Abstract] |
| [Author] | Author name | "Zhang F"[Author] |
| [Journal] | Journal name | "Nature"[Journal] |
| [Publication Type] | Article type filter | "Review"[Publication Type] |
| [Date - Publication] | Publication date range | "2020/01/01"[Date - Publication]:"2024/12/31"[Date - Publication] |
| [MeSH Major Topic] | MeSH term as major focus of the article | "CRISPR-Cas Systems"[MeSH Major Topic] |
Boolean Operators
Boolean operators control how search terms combine:
# AND: All terms must be present -- narrows results
results = query_pubmed("CRISPR AND cancer AND therapy")
# OR: Any term can be present -- broadens results (use for synonyms)
results = query_pubmed("(tumor OR tumour OR neoplasm) AND immunotherapy")
# NOT: Exclude terms -- use sparingly to avoid losing relevant papers
results = query_pubmed("cancer immunotherapy NOT review")
Use parentheses to group OR terms together before combining with AND.
arXiv Subject Categories
arXiv organizes preprints by subject category. Biology-related categories include:
| Category | Description |
|----------|-------------|
| q-bio.BM | Biomolecules |
| q-bio.CB | Cell Behavior |
| q-bio.GN | Genomics |
| q-bio.MN | Molecular Networks |
| q-bio.NC | Neurons and Cognition |
| q-bio.QM | Quantitative Methods |
| cs.AI | Artificial Intelligence |
| cs.LG | Machine Learning |
Decision Framework
Use this tree to determine which search tier and database to start with:
What type of question are you answering?
├── Clinical / biomedical question
│ ├── Specific drug or treatment → Tier 1: PubMed with PICO query
│ ├── Disease mechanism → Tier 1: PubMed with MeSH terms
│ └── Clinical trial evidence → Tier 1: PubMed filtered by Publication Type
├── Computational / quantitative methods
│ ├── ML model or algorithm → Tier 1: arXiv (cs.LG, cs.AI)
│ ├── Computational biology method → Tier 1: arXiv (q-bio.*) + PubMed
│ └── Software tool or pipeline → Tier 2: AI-assisted web search
├── Broad research landscape
│ ├── Current state of a field → Tier 2: AI-assisted web search
│ ├── Recent developments (last 6 months) → Tier 2: AI-assisted web search
│ └── Cross-disciplinary question → Tier 1: Google Scholar + Tier 2
├── Specific paper or data
│ ├── Known paper details → Tier 1: any database by title/author/DOI
│ ├── Methodology or protocol → Tier 3: full-text extraction
│ └── Supplementary data → Tier 3: DOI-based supplementary fetch
└── Protocols / reagents
├── Lab protocol → Tier 2: web search for protocols.io, etc.
└── Validated reagents → Tier 2: AI-assisted web search
| Scenario | Recommended Tier and Database | Rationale | |----------|-------------------------------|-----------| | Systematic review of clinical evidence | Tier 1: PubMed with MeSH + publication type filters | Reproducible, documented search strategy required | | Finding a preprint on a new ML method | Tier 1: arXiv with category and keyword search | Preprints appear on arXiv before journals | | Understanding the research landscape | Tier 2: AI-assisted web search | Requires synthesis across many sources | | Extracting a specific protocol from a paper | Tier 3: PDF content extraction | Need full-text access to methods section | | Finding papers across disciplines | Tier 1: Google Scholar | Broadest coverage across fields | | Identifying key researchers in a niche area | Tier 2: AI-assisted web search | Requires contextual synthesis | | Downloading supplementary data tables | Tier 3: DOI-based supplementary fetch | Direct access to supplementary files |
Best Practices
-
Use controlled vocabulary (MeSH) for PubMed searches: Free-text searches miss papers that use different terminology. MeSH terms map synonyms to a single concept, improving recall without sacrificing precision.
# Free text misses synonyms query_pubmed("heart attack treatment") # MeSH captures all synonyms query_pubmed('"Myocardial Infarction"[MeSH] AND "Drug Therapy"[MeSH]') -
Include synonyms and alternative terms with OR: Scientific concepts often have multiple names (e.g., tumor/tumour/neoplasm). Group synonyms with OR inside parentheses to avoid missing relevant papers.
query_pubmed("(myocardial infarction OR heart attack) AND (treatment OR therapy)") -
Use phrase searching for multi-word concepts: Quoting exact phrases prevents the search engine from splitting terms and matching them independently.
query_pubmed('"single cell RNA sequencing" AND methods') -
Filter by publication type when seeking specific evidence: Clinical trials, systematic reviews, and meta-analyses each answer different questions. Use
[Publication Type]to target the evidence level you need.query_pubmed("COVID-19 vaccine efficacy AND clinical trial[Publication Type]") -
Start broad, then narrow iteratively: Begin with core concepts (2-3 terms) and review initial results. Add specificity based on what you find -- more terms, date ranges, field tags, or publication types.
# Step 1: Broad results = query_pubmed("CRISPR base editing iPSC", max_papers=20) # Step 2: Add MeSH and specificity results = query_pubmed( '"CRISPR-Cas Systems"[MeSH] AND "base editing" AND "induced pluripotent stem cells" AND efficiency', max_papers=20 ) # Step 3: Filter by date results = query_pubmed( '"CRISPR-Cas Systems"[MeSH] AND "base editing" AND "induced pluripotent stem cells" AND efficiency AND ("2022"[Date - Publication]:"2024"[Date - Publication])', max_papers=20 ) -
Cross-reference multiple databases: No single database covers all literature. Use PubMed for biomedical content, arXiv for computational preprints, and Google Scholar for cross-disciplinary coverage.
-
Assess result quality systematically: Evaluate papers for source reliability (peer-reviewed journal), author credentials, recency, study design appropriateness, sample size adequacy, reproducibility, declared conflicts of interest, and citation count.
Common Pitfalls
-
Overly long and specific queries: Packing too many terms into a single query causes missed results because all terms must match simultaneously.
- How to avoid: Limit queries to core concepts (3-5 terms). Run separate searches for sub-topics and combine results manually.
# Too specific -- misses relevant papers query_pubmed("CRISPR Cas9 gene editing HEK293T cells 2024 efficiency optimization delivery") # Better -- core concepts only query_pubmed("CRISPR Cas9 gene editing optimization efficiency") -
Relying on a single database: PubMed has biomedical focus, arXiv covers preprints, Google Scholar spans disciplines. Using only one database guarantees blind spots.
- How to avoid: Always search at least two databases. For computational biology, combine PubMed and arXiv. For cross-disciplinary topics, include Google Scholar.
-
Ignoring publication dates: Scientific knowledge evolves rapidly. Foundational papers remain relevant, but methods and clinical evidence may be superseded.
- How to avoid: Check publication dates in all results. For methods papers, prefer the last 3-5 years. For foundational concepts, older papers are acceptable but verify with recent reviews.
-
Skipping title and abstract review before deep-diving: Not all search results that match keywords are actually relevant. Downloading and reading full texts without screening wastes time.
- How to avoid: Always screen titles and abstracts first. Only extract full text (Tier 3) for papers that pass screening.
-
Using NOT operators too aggressively: The NOT operator can inadvertently exclude relevant papers that mention the excluded term in a different context.
- How to avoid: Use NOT sparingly
Truncated for display — read the full file on GitHub.
Related Skills
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
ui-ux-pro-max
130.2kUI/UX design intelligence for web, mobile, and desktop. This skill should be used when designing, building, reviewing, or fixing interfaces, including pages, components, design systems, accessibility, interaction, responsive layout, typography, color, charts, and stack-specific UI implementation.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
