SkillAgentSearch skills...

literature-review-agent

Step 3 of the PaperOrchestra pipeline (arXiv:2604.05018). Execute the literature search strategy from outline.json — discover candidate papers via web search, verify them through Semantic Scholar (Levenshtein > 70 fuzzy title match, temporal cutoff, dedup by paperId), cross-corroborate against Cross…

Install / Use

npx skills add Ar9av/PaperOrchestra --skill literature-review-agent

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

92/100

Category

Automation

Supported Platforms

Universal

Our assessment of literature-review-agent

literature-review-agent scores 92/100 on our quality scale, 957th of 2,855 Automation skills we index (top 34%).

Its SKILL.md is 20 KB long, well organised into 39 sections with 19 code examples: a thorough specification that gives an agent plenty to work with.

It has 664 GitHub stars, a meaningful sign that others use it.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
12/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 12 days ago, so literature-review-agent is actively maintained.
  • No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
  • Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.

Automated pattern scan on 2026-10-04. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.

literature-review-agent compared with similar skills

All 4 of these similar skills score higher than literature-review-agent; compare them before choosing.

SkillScoreStarsUpdatedFormat
literature-review-agent (this skill)by Ar9av9266412d agoSKILL.md
Agent-Reachby Panniantong10090.1k18d agoCLAUDE.md
Scraplingby D4Vinci10085.6k1d agoMCP Server
rufloby ruvnet10073.8ktodayMCP Server
algorithmic-artby anthropics100177.9k11d agoSKILL.md

Frequently asked questions

How do I install literature-review-agent?
Run npx skills add Ar9av/PaperOrchestra --skill literature-review-agent. The install tabs above show the steps for each supported agent.
Which AI agents does literature-review-agent work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is literature-review-agent safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is literature-review-agent still maintained?
The repository was last updated 12 days ago, so literature-review-agent is actively maintained.

name: literature-review-agent description: Step 3 of the PaperOrchestra pipeline (arXiv:2604.05018). Execute the literature search strategy from outline.json — discover candidate papers via web search, verify them through Semantic Scholar (Levenshtein > 70 fuzzy title match, temporal cutoff, dedup by paperId), cross-corroborate against Crossref + OpenAlex to flag hallucinated citations, build a BibTeX file, and draft Introduction + Related Work using ≥90% of the verified pool. Runs in parallel with the plotting-agent. TRIGGER when the orchestrator delegates Step 3 or when the user asks to "find citations for my paper", "draft the related work", or "build the bibliography".

Literature Review Agent (Step 3)

Faithful implementation of the Hybrid Literature Agent from PaperOrchestra (Song et al., 2026, arXiv:2604.05018, §4 Step 3, App. D.3, App. F.1 p.46).

Cost: ~20–30 LLM calls. This is one of the two longest steps (the other is plotting). Wall-time floor is set by Semantic Scholar's 1 QPS verification limit.

Inputs

  • workspace/outline.json — specifically intro_related_work_plan with the Introduction search directions and the 2-4 Related Work methodology clusters
  • workspace/inputs/conference_guidelines.md — used to derive cutoff_date
  • workspace/inputs/idea.md, workspace/inputs/experimental_log.md — for framing the Intro and grounding the Related Work positioning

Outputs

  • workspace/citation_pool.json — verified Semantic Scholar metadata for every paper that survived verification
  • workspace/refs.bib — BibTeX file generated from the verified pool
  • workspace/drafts/intro_relwork.tex — drafted Introduction and Related Work sections, written into the template, with the rest of the template preserved verbatim

Two-phase pipeline (App. D.3)

PHASE 1 — Parallel Candidate Discovery
   For each search direction in introduction_strategy.search_directions:
   For each limitation_search_query in each related_work cluster:
     - Use the host's web search tool to discover up to ~10 candidate papers.
     - Run up to 10 discovery queries in parallel (host-permitting).
     - Collect (title, snippet, url) tuples — no verification yet.
   → PRE-DEDUP before Phase 2 (see Step 1.5 below)

PHASE 2 — Sequential Citation Verification (1 QPS, with cache)
   For each candidate (after pre-dedup), sequentially:
     0. Check s2_cache.json first (scripts/s2_cache.py --check).
        If HIT: use cached response, skip live S2 call. No throttle needed.
        If MISS: proceed with live request below.
     1. Query Semantic Scholar by title:
          GET https://api.semanticscholar.org/graph/v1/paper/search?query=<title>
              &fields=title,abstract,year,authors,venue,externalIds&limit=5
        (Public endpoint, no key. Throttle to 1 QPS for live requests only.)
     2. Store the S2 response in cache: s2_cache.py --store.
     3. Pick the top hit. Check Levenshtein title ratio against the original
        candidate title. If ratio < 70: discard.
     4. Bonus: if year and venue exactly align with hints, add a +5 point
        match-quality bonus.
     5. Require: abstract is non-empty.
     6. Require: paper.year (or month if known) strictly predates cutoff_date.
        Months default to day-1: e.g., "October 2024" → 2024-10-01.
     7. If all checks pass, add to verified pool.
   After all candidates are verified, dedup by Semantic Scholar paperId.

The host agent does the LLM/web work; the deterministic helpers in scripts/ do the math.

Step-by-step

0. Derive cutoff_date

Parse conference_guidelines.md for the submission deadline. The paper aligns research cutoff with venue submission deadline (App. D.1):

| Venue | Cutoff | |---|---| | CVPR 2025 | Nov 2024 | | ICLR 2025 | Oct 2024 | | Other | One month before the stated submission deadline |

Encode as YYYY-MM-DD. Months default to day-1 (e.g., 2024-10-01).

1. Phase 1: Parallel Candidate Discovery

From outline.json:

  • All introduction_strategy.search_directions (3-5 queries)
  • For each cluster in related_work_strategy.subsections:
    • The cluster's sota_investigation_mission becomes a search query
    • All limitation_search_queries (1-3 each)

For each query, use your host's web search tool (e.g., WebSearch in Claude Code, @web in Cursor, the search tool in Antigravity). Collect the top ~10 candidates per query: title, abstract snippet, source URL.

If your host supports parallel sub-tasks, fire up to 10 concurrent search queries. If not, run sequentially — slower but functionally equivalent.

Optional: Exa as a Phase 1 backend

If your host has no native web search, OR you want a research-paper-focused backend with better signal-to-noise, you can use Exa via the bundled scripts/exa_search.py helper. It is opt-in and reads EXA_API_KEY from the environment — the repo never commits a key.

export EXA_API_KEY="your-key-here"   # get one at https://dashboard.exa.ai/
python skills/literature-review-agent/scripts/exa_search.py \
    --query "Sparse attention long context transformers" \
    --num-results 15 \
    --discovered-for "related_work[2.1]"

Output is a normalized candidate list ready to merge into raw_candidates.json. Phase 2 verification (Semantic Scholar fuzzy match, cutoff, dedup) is unchanged. See references/exa-search-cookbook.md for the full recipe, query patterns, cost estimates, and security notes.

Optional: Tavily as a Phase 1 backend

If your host has no native web search, OR you want an LLM-optimized search backend with high relevance scoring, you can use Tavily via the bundled scripts/tavily_search.py helper. It is opt-in and reads TAVILY_API_KEY from the environment — the repo never commits a key.

export TAVILY_API_KEY="tvly…[redacted]"   # get one at https://app.tavily.com
python skills/literature-review-agent/scripts/tavily_search.py \
    --query "Sparse attention long context transformers" \
    --num-results 15 \
    --academic \
    --discovered-for "related_work[2.1]"

Output is a normalized candidate list ready to merge into raw_candidates.json. Phase 2 verification (Semantic Scholar fuzzy match, cutoff, dedup) is unchanged. See references/tavily-search-cookbook.md for the full recipe, query patterns, cost estimates, and security notes.

Combine all discovered candidates into a single working list. Tag each with the originating query ID so you can later attribute it to "intro" vs "related_work[i]".

1.5. Pre-dedup before Phase 2

Always run this before starting Phase 2. Multiple search queries routinely return the same papers (e.g., "Attention is All You Need" appears in almost every NLP discovery query). Verifying duplicates wastes 30-40% of S2 quota at 1 QPS.

python skills/literature-review-agent/scripts/pre_dedup_candidates.py \
    --in workspace/raw_candidates.json \
    --out workspace/deduped_candidates.json
# Prints: "150 candidates → 97 unique (53 duplicates removed)"

Use workspace/deduped_candidates.json as input to Phase 2.

2. Phase 2: Sequential Verification via Semantic Scholar (with cache)

For each candidate in deduped_candidates.json, in sequential order:

Step A — check cache first (no S2 call, no throttle needed):

python skills/literature-review-agent/scripts/s2_cache.py \
    --cache workspace/cache/s2_cache.json \
    --check "<candidate title>"
# exit 0 + prints JSON → use cached response, skip Step B
# exit 1 → proceed to Step B

Step B — live S2 request (cache MISS only, throttle to 1 QPS):

Preferred: use the bundled scripts/s2_search.py helper — it handles auth, retries, and 429 back-off automatically:

python skills/literature-review-agent/scripts/s2_search.py \
    --query "<URL-decoded candidate title>" --limit 5
# If SEMANTIC_SCHOLAR_API_KEY is set the key is forwarded automatically.
# If not, the public unauthenticated endpoint is used (≤1 QPS, still works).

Check whether the key is configured before starting Phase 2:

python skills/literature-review-agent/scripts/s2_search.py --check-key

Fallback: if you prefer your host's URL fetch tool, GET:

https://api.semanticscholar.org/graph/v1/paper/search?query=<URL-encoded title>&limit=5&fields=title,abstract,year,authors,venue,externalIds

Add header x-api-key: <SEMANTIC_SCHOLAR_API_KEY> if the env var is set. Be polite: ≤1 request per second for live requests. Cache hits are free.

Step C — store in cache (after every successful live request):

python skills/literature-review-agent/scripts/s2_cache.py \
    --cache workspace/cache/s2_cache.json \
    --store "<candidate title>" \
    --response '<full S2 JSON response>'

For the top hit:

python skills/literature-review-agent/scripts/levenshtein_match.py \
    --candidate "Original candidate title" \
    --found "S2 returned title"
# prints integer 0-100. Discard if < 70.

Then check the temporal cutoff:

python skills/literature-review-agent/scripts/check_cutoff.py \
    --paper-year 2024 \
    --paper-month 9 \
    --cutoff 2024-10-01
# exit 0 if strictly predates, exit 1 if not

If both checks pass AND the abstract is non-empty, append the paper's full S2 metadata to the verified pool.

3. Dedup and assemble the pool

After all candidates are verified:

python skills/literature-review-agent/scripts/dedupe_by_id.py \
    --in raw_pool.json \
    --out workspace/citation_pool.json

The dedupe script keys on paperId (Semantic Scholar's internal unique ID), falling back to externalIds.DOI, then externalIds.ArXiv, then a normalized title.

The script also computes and writes min_cite_paper_count = floor(0.9 * len(papers)) — the minimum number of papers the writing step must cite (the paper's ≥90% integration rule, App. D.3).

Immediately after dedupe_by_id.py, validate and auto-fix the pool schema:

python skills/literature-review-agent/scripts/validate_pool.py \
    --pool workspace/citation_pool.json --fix
# Catches and fixes authors-as-strings, reports missing required fields.
# Must pass before proceeding to Step 4.

3.5. Cross-index verification (Crossref + OpenAlex)

Semantic Scholar is one index and can return a plausible record for a paper that does not exist, or attach wrong metadata. Re-check every S2-verified paper against two independent indices before building the bibliography — this is the practical defense against hallucinated citations leaking in.

# Optional but recommended: a polite-pool email gives faster, more reliable
# service. The repo never commits an address.
export PAPER_ORCHESTRA_MAILTO="you@example.com"

python skills/literature-review-agent/scripts/cross_verify.py \
    --pool workspace/citation_pool.json --inplace
# Annotates each paper with a `cross_verification` field and writes
# workspace/cross_verification_report.json.
# exit 0 = all corroborated; exit 1 = WARN (something flagged or an index
# was unreachable); exit 2 = usage error.

This is a WARN gate, not a hard gate (like validate_consistency.py): it flags suspicious citations but does not block the pipeline or delete anything. Review the low and conflict tiers in the report:

  • high — corroborated by ≥1 external index → keep.
  • medium — corroborated but year disagrees → keep, spot-check the year.
  • low — not found in Crossref or OpenAlex → review by hand. Note that arXiv-only preprints (no DOI) are a common benign cause; low means "could not corroborate," not "fabricated." S2 already confirmed it exists.
  • conflict — pool DOI disagrees with the external DOI → likely wrong record.

Drop only the entries you genuinely cannot corroborate, then re-run dedupe_by_id.py onward. If both indices are unreachable (offline), the script degrades gracefully and the pipeline c

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars664
CategoryAutomation
Updated12d ago
Forks92

Languages

Python

Trust signals

88/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 medium