SkillAgentSearch skills...

arxiv-mcp-server

Search arXiv, fetch paper metadata, and read full-text content via MCP. STDIO & Streamable HTTP.

Install / Use

claude mcp add cyanheads -- npx -y github:cyanheads/arxiv-mcp-server

If the server publishes to npm under a different name, use that package instead — check the repo README.

About this skill
🔌

MCP Server

Model Context Protocol server

Quality Score

80/100

Category

Marketing

Supported Platforms

Claude Code
Claude Desktop
<div align="center"> <h1>@cyanheads/arxiv-mcp-server</h1> <p><b>Search arXiv, fetch paper metadata, and read full-text content via MCP. STDIO or Streamable HTTP.</b> <div>4 Tools • 2 Resources</div> </p> </div> <div align="center">

Version License Docker MCP SDK npm TypeScript

</div> <div align="center">

Install in Claude Desktop Install in Cursor Install in VS Code

Framework

</div> <div align="center">

Public Hosted Server: https://arxiv.caseyjhand.com/mcp

</div>

Tools

Four tools for searching and reading arXiv papers:

| Tool Name | Description | |:----------|:------------| | arxiv_search | Search arXiv papers by query with category and sort filters. | | arxiv_get_metadata | Get full metadata for one or more arXiv papers by ID. | | arxiv_read_paper | Fetch the full text content of an arXiv paper from its HTML rendering, or from the PDF when no render exists. | | arxiv_list_categories | List arXiv category taxonomy, optionally filtered by group. |

arxiv_search

Search for papers using free-text queries with field prefixes and boolean operators.

  • Field prefixes: ti: (title), au: (author), abs: (abstract), cat: (category), all: (all fields)
  • Boolean operators: AND, OR, ANDNOT
  • Optional category filter, sorting (relevance, submitted, updated), and pagination
  • Category accepts a leaf code (cs.CL) or a whole archive (astro-ph, cs, math) — a bare archive covers its subject classes plus the legacy flat papers filed before it was subdivided
  • submitted_from / submitted_to bound the submission date (inclusive, UTC YYYY-MM-DD). Consecutive windows cover the matches with no gap — a paper submitted exactly at a midnight seam falls in both, so de-duplicate by ID — which is how to reach results past the 10,000 pagination ceiling
  • Echoes back the query as actually searched, with every filter folded in — replaying it reproduces the same result set
  • Returns up to 50 results per request with full metadata including abstract

arxiv_get_metadata

Fetch full metadata for one or more papers by known arXiv ID.

  • Batch fetch up to 10 papers in a single request
  • Accepts both versioned (2401.12345v2) and unversioned (2401.12345) IDs
  • Legacy ID format supported (hep-th/9901001)
  • Reports not-found IDs separately from found papers

arxiv_read_paper

Read the full body of an arXiv paper.

  • Tries native arXiv HTML first, then ar5iv, then text extracted from the PDF — the source field reports which one answered
  • Strips HTML head/boilerplate and collapses MathML to dollar-delimited LaTeX ($…$ inline, $$…$$ block) so the character budget targets paper content
  • Returns raw HTML — no parsing or extraction; the LLM interprets content directly. PDF-extracted bodies are plain text: prose is reliable, but math, tables, and heading structure flatten
  • max_characters defaults to 100,000; pass null for the whole paper in one call. Raw HTML can be 500KB-3MB+ for math-heavy papers, which is more than most clients accept in a single tool result — page with start instead

arxiv_list_categories

List arXiv category codes and names for discovery.

  • ~155 categories across 8 top-level groups (cs, math, physics, q-bio, q-fin, stat, eess, econ)
  • Optional group filter to narrow results
  • Static data — always succeeds

Resources

| URI Pattern | Description | |:------------|:------------| | arxiv://paper/{paperId} | Paper metadata by arXiv ID. Percent-encode a legacy ID's slash — arxiv://paper/hep-th%2F9901001. | | arxiv://categories | Full arXiv category taxonomy. |

Features

Built on @cyanheads/mcp-ts-core:

  • Declarative tool definitions — single file per tool, framework handles registration and validation
  • Unified error handling across all tools
  • Pluggable auth (none, jwt, oauth)
  • Structured logging with optional OpenTelemetry tracing
  • Runs locally (stdio/HTTP) from the same codebase

arXiv-specific:

  • Read-only, no authentication required — arXiv API is free, metadata is CC0
  • Rate-limited request queue enforcing arXiv's 3-second crawl delay
  • Adaptive cooldown on rate-limit (5s → 10s → 20s → 30s), honors Retry-After
  • Retry with exponential backoff for transient failures
  • Content fallback chain: native arXiv HTML → ar5iv → PDF text extraction (both HTML renders run LaTeXML, so they tend to fail together; the PDF is the artifact every paper has, and it also covers an ar5iv outage rather than letting one fail the read)
  • Full arXiv category taxonomy embedded as static data
  • Optional local OAI-PMH metadata mirror (SQLite + FTS5) — opt-in, eliminates rate-limit exposure for arxiv_search and arxiv_get_metadata. See Optional: Local Mirror.

Getting Started

Public Hosted Instance

A public instance is available at https://arxiv.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:

{
  "mcpServers": {
    "arxiv-mcp-server": {
      "type": "streamable-http",
      "url": "https://arxiv.caseyjhand.com/mcp"
    }
  }
}

Self-Hosted / Local

Add to your MCP client config (e.g., claude_desktop_config.json):

{
  "mcpServers": {
    "arxiv-mcp-server": {
      "type": "stdio",
      "command": "bunx",
      "args": ["@cyanheads/arxiv-mcp-server@latest"]
    }
  }
}

Prerequisites

Installation

  1. Clone the repository:
git clone https://github.com/cyanheads/arxiv-mcp-server.git
  1. Navigate into the directory:
cd arxiv-mcp-server
  1. Install dependencies:
bun install

Configuration

All configuration is optional — the server works out of the box with sensible defaults.

| Variable | Description | Default | |:---------|:------------|:--------| | ARXIV_API_BASE_URL | arXiv API base URL. | https://export.arxiv.org/api | | ARXIV_REQUEST_DELAY_MS | Minimum delay between arXiv API requests (ms). | 3000 | | ARXIV_CONTENT_TIMEOUT_MS | Timeout for paper body fetches — HTML renders and PDF downloads (ms). | 30000 | | ARXIV_API_TIMEOUT_MS | Timeout for API search/metadata requests (ms). | 15000 | | ARXIV_MIRROR_ENABLED | Enable local OAI-PMH metadata mirror for search and metadata. | false | | ARXIV_MIRROR_PATH | SQLite path for the mirror. | ./data/arxiv-mirror.db | | ARXIV_MIRROR_REFRESH_CRON | UTC cron expression for in-process daily refresh (HTTP mode only). | unset | | ARXIV_MIRROR_FALLBACK_LIVE | Fall through to live API on local ID-lookup miss. | true | | ARXIV_MIRROR_RECENT_DAYS_LIVE | Route sortBy=submitted descending queries within this window to the live API. | 2 | | ARXIV_MIRROR_OAI_BASE_URL | arXiv OAI-PMH endpoint base URL. | https://oaipmh.arxiv.org/oai | | ARXIV_MIRROR_OAI_REQUEST_DELAY_MS | Minimum delay between OAI-PMH requests (ms). | 3000 | | ARXIV_MIRROR_REFRESH_TIMEOUT_MS | Abort budget for one scheduled refresh subprocess (ms). | 7200000 | | MCP_TRANSPORT_TYPE | Transport: stdio or http. | stdio | | MCP_HTTP_PORT | Port for HTTP server. | 3010 | | MCP_AUTH_MODE | Auth mode: none, jwt, or oauth. | none | | MCP_LOG_LEVEL | Log level (RFC 5424). | info |

Running the Server

Local Development

  • Build and run:

    bun run build
    bun run start:http   # or start:stdio
    
  • Run checks and tests:

    bun run devcheck     # Lint, format, typecheck, audit
    bun run test         # Vitest
    

Optional: Local Mirror

For self-hosted deployments behind a single egress IP, arXiv's ~3-second per-IP crawl delay serializes concurrent users. An optional local mirror eliminates rate-limit exposure for arxiv_search and arxiv_get_metadata by serving from a SQLite + FTS5 store harvested via OAI-PMH. arxiv_read_paper continues to use the live API — full-content harvest is forbidden by arXiv's data policy.

Disabled by default. To enable:

# 1. Cold-start harvest (~4.4h sequential, resumable from checkpoint). One-time per installation.
bun run mirror:init

# 2. Enable the mirror.
export ARXIV_MIRROR_ENABLED=true

# 3. Start the server — reads switch to the mirror once the harvest completes.
bun run start:http

Daily incremental refresh (small delta; duration depends on arXiv's OAI-PMH page pacing) via:

bun run mirror:refresh   # wire to cron / systemd timer / launchd, OR
                         # set ARXIV_MIRROR_REFRESH_CRON to schedule it in HTTP mode (spawned as a child process)
bun run mirror:verify    # schema version + PRAGMA integrity_check / quick_check

Schema upgrades. The mirror records a schema version and migrates itself in place the first time a newer server opens it — never a re-harvest, and never a separate operator step. The upgrade that added comment and journal_ref to the full-text index (#37) rebuilds that index from the rows already stored, so co: and jr: searches resolve against a mirror harvested before it. The rebuild runs at startup, before the store answers its first read, and logs mirror migration v2→v3 (fts rebuild) progress lines throughout — on a full-corpus mirror, expect the first start after the upgrade to take noticeably longer than usual. An interrupted rebuild is repeated on the next open rather than left half-applied. bun run mirror:verify prints the schema version the file carries and exits non-zero if a migration never completed.

Behavior notes. Ranking divergence: FTS5 BM25 differs from arXiv's internal ranking, so sortBy=relevance against the mirror returns a different top-K than the live API. Queries sorted by submitted descending within ARXIV_MIRROR_RECENT_DAYS_LIVE days route to the live API to cover the nightly-update gap. Refresh resilience: after the initial cold harvest completes, an in-progress or failed daily refresh keeps serving the existing dataset from the mirror — arxiv_search and arxiv_get_metadata don't drop to the live API during the refresh window (#21). The scheduled HTTP-mode refresh runs in a child process, so the harvest's synch

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars3
CategoryMarketing
Updated20d ago
Forks0

Languages

TypeScript

Security Score

87/100

Audited on Jul 27, 2026

2 low