docshelf-mcp
MCP server that turns a folder of PDFs and Markdown into an AI-friendly document shelf. Convert, split by chapter, auto-index — Claude/ChatGPT answers from a 5 KB INDEX over raw URLs.
Install / Use
claude mcp add ignatenkofi -- npx -y github:ignatenkofi/docshelf-mcpIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AI & Machine LearningSupported Platforms
Our assessment of docshelf-mcp
docshelf-mcp scores 83/100 on our quality scale, 641st of 961 AI & Machine Learning skills we index.
Its MCP Server is 14 KB long, well organised into 28 sections with 13 code examples: a thorough specification that gives an agent plenty to work with.
It has 3 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated 2 days ago, so docshelf-mcp is actively maintained.
- Our last check on 2026-09-06 found the source still online.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 92/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
docshelf-mcp compared with similar skills
All 4 of these similar skills score higher than docshelf-mcp; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| docshelf-mcp (this skill)by ignatenkofi | 83 | 3 | 2d ago | MCP Server |
| claude-memby thedotmack | 100 | 96.4k | today | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 91.6k | 20d ago | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.3k | 3d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.4k | today | CLAUDE.md |
Frequently asked questions
- How do I install docshelf-mcp?
- Run
claude mcp add ignatenkofi -- npx -y github:ignatenkofi/docshelf-mcp. The install tabs above show the steps for each supported agent. - Which AI agents does docshelf-mcp work with?
- It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
- Is docshelf-mcp safe to use?
- It is MIT-licensed and scores 92/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is docshelf-mcp still maintained?
- The repository was last updated 2 days ago, so docshelf-mcp is actively maintained.
Skill content
View source on GitHubdocshelf-mcp
Put your manuals on a shelf, hand the AI the index.
📖 Docs & landing page: https://ignatenkofi.github.io/docshelf-mcp/
_ _ _ __
__| | ___ ___ ___| |__ ___| |/ _|
/ _` |/ _ \ / __/ __| '_ \ / _ \ | |_
| (_| | (_) | (__\__ \ | | | __/ | _|
\__,_|\___/ \___|___/_| |_|\___|_|_|
MCP server for AI-friendly doc shelves
An MCP server that turns a folder of PDFs and Markdown into a chat-project-friendly document collection: AI agents see a single INDEX.md and pull individual sections by raw GitHub URL on demand — instead of choking on a 4 MB datasheet.
Why?
You have 30 hardware manuals, or 200 cooking recipes, or a stack of research PDFs.
You want Claude / ChatGPT / whatever to be able to answer questions across them — but:
- ❌ You can't dump 80 MB of PDFs into a chat project. It won't fit, and you'd burn the context window even if it did.
- ❌ You can manually copy-paste the relevant pages, but only after you remember which manual mentioned the thing you need.
- ❌ Long files mean retrieval is wasteful — the model loads the whole RouterOS guide just to answer a question about VLANs.
docshelf-mcp solves it like this:
- You drop a PDF onto the shelf.
- The shelf converts it to Markdown, splits big files chapter-by-chapter, and regenerates a navigation
INDEX.md. - You commit and push to a public GitHub repo.
- Add only
INDEX.mdto your Claude project. When the model needs a section, it fetches it viaraw.githubusercontent.com.
Result: a 5 KB index pointing at a 50 MB collection. The model reads exactly the chapter it needs.
📦 Install
From PyPI (once the first tagged release is published):
# uv (recommended)
uv pip install docshelf-mcp
# or plain pip
pip install docshelf-mcp
Or straight from main (always-latest, no PyPI required):
pip install "git+https://github.com/ignatenkofi/docshelf-mcp"
Optional high-quality PDF engine (pulls ~2 GB of PyTorch — only if you need it):
pip install "docshelf-mcp[high-quality]"
Optional input formats beyond PDF/Markdown — DOCX, HTML, EPUB (lightweight):
pip install "docshelf-mcp[formats]" # or [docx] / [html] / [epub]
📋 Project Prompt
Drop this into the Custom Instructions of any Claude project that consumes
a docshelf-style INDEX.md:
This project uses the docshelf pattern.
INDEX.mdis the entry point. When answering: read INDEX → fetch ONLY the needed section file via its GitHub raw URL (use WebFetch / fetch / curl). Don't load full source files into context. For large manuals split into chapters, follow INDEX → chapter SUBINDEX → section file.
Medium (~150 words) and full (~400 words) versions, plus how-to snippets for
Claude Code, Claude Desktop, and the Anthropic API, live in
docs/PROJECT_PROMPT.md.
Quickstart (Python library)
from docshelf_mcp import Shelf
shelf = Shelf("~/Documents/my-homelab-docs").init(
name="My HomeLab Docs",
remote="https://github.com/me/my-homelab-docs",
default_categories=["routers", "switches", "psu", "motherboards"],
)
shelf.add_document(
"~/Downloads/MIKROTIK_RouterOS.pdf",
category="routers",
title="Mikrotik RouterOS — full manual",
description="Official RouterOS reference, split by chapter.",
)
# → docs/routers/mikrotik-routeros-full-manual.md + docs/routers/.../001-..md, 002-..md, ...
# → INDEX.md is regenerated automatically.
Then in the shelf directory: git add . && git commit -m "docs: add RouterOS" && git push.
In your Claude project, attach only INDEX.md. Done.
Quickstart (MCP server)
1. Add to Claude Desktop
Edit ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%/Claude/claude_desktop_config.json (Windows):
{
"mcpServers": {
"docshelf": {
"command": "docshelf-mcp",
"env": {
"DOCSHELF_ROOT": "/Users/me/Documents/my-homelab-docs"
}
}
}
}
Restart Claude Desktop. You now have eleven new tools available:
| Tool | What it does |
|---|---|
| docshelf_init_shelf | Bootstrap a new shelf directory. |
| docshelf_add_document | Add a file (MD/PDF/DOCX/HTML/EPUB). Converts, splits, re-indexes. |
| docshelf_add_directory | Add every supported file (MD/PDF/DOCX/HTML/EPUB) in a folder in one call. Re-indexes once. |
| docshelf_read_document | Read a document/section's content over MCP (works on private shelves). |
| docshelf_remove_document | Remove a document, its sections, and metadata. Re-indexes. |
| docshelf_rename_document | Retitle / recategorize a document (moves file, sections, meta) — no re-conversion. |
| docshelf_rebuild_index | Regenerate INDEX.md from disk. |
| docshelf_doctor | Check shelf integrity; optionally auto-fix safe drift. |
| docshelf_search | Plain-text search across the shelf, with raw URLs. |
| docshelf_list_documents | List documents by category. |
| docshelf_convert_pdf | Standalone PDF → Markdown (no shelf). |
The shelf files are also exposed as read-only MCP resources, so a client can browse and attach them natively — see MCP Resources below.
2. Add to Claude Code
claude mcp add docshelf -- docshelf-mcp
# Optional: set the default shelf
claude mcp add docshelf --env DOCSHELF_ROOT=/path/to/shelf -- docshelf-mcp
3. Test from the command line
# Sanity check — should print the server version then wait on stdin
docshelf-mcp
MCP Resources
Alongside the tools, every shelf file is exposed as a read-only MCP resource, so an MCP client (Claude Desktop, Claude Code, …) can browse and attach shelf content natively — no tool call required.
- Scheme:
docshelf:///<relative-path>, e.g.docshelf:///INDEX.mdordocshelf:///docs/routers/mikrotik/003-firewall.md. - What's exposed:
INDEX.mdplus every document and every split section underdocs/— one resource each. A split document exposes both its whole-file parent and its individual section files. - Size cap: a resource read is capped at 1 MB (1,000,000 bytes). A larger file is truncated at a UTF-8 character boundary and ends with a notice pointing at the
docshelf_read_documenttool, which pages the rest. - Freshness: content is read from disk on every access, and the resource set is re-synced when the server starts and after each mutating tool call (
add_document,add_directory,remove_document,rename_document,rebuild_index,init_shelf) — so newly added files appear and removed ones drop out. Reads are confined to the shelf root.
Resources are only registered for an initialized shelf (one that has a .docshelf.json); a non-shelf DOCSHELF_ROOT simply exposes none.
The shelf layout
my-shelf/
├── .docshelf.json ← shelf metadata: name, remote, category order
├── INDEX.md ← auto-generated navigation (your chat-project file)
├── .gitignore
└── docs/
├── routers/
│ ├── .meta.json ← per-document title/description overrides
│ ├── mikrotik-routeros.md (full document, lightly cleaned)
│ └── mikrotik-routeros/ (auto-split sections)
│ ├── SUBINDEX.md (per-document navigation page)
│ ├── 001-overview.md
│ ├── 002-bridging.md
│ └── 003-firewall.md
└── switches/
└── cudy-gs1010pe.md
Everything in docs/ is committed; everything is fetchable via raw URL once you push to GitHub.
How splitting works
A document is split when both conditions hold:
- UTF-8 size > 50 KB (configurable via
.docshelf.json:split_threshold_bytes). - The document has at least two
##(H2) headings.
The splitter:
- Cleans PDF-extraction noise (collapses runaway blank lines, demotes CLI dumps mistaken for H1s).
- Slices on H2 boundaries.
- Names files
NNN-<slug>.mdso they sort naturally and survive title changes. - Wipes the previous split directory before regenerating — fully idempotent.
- Writes a
SUBINDEX.mdnavigation page into the split directory (title, description, per-section links) — regenerated on everyrebuild_index.
In INDEX.md, split documents with up to 10 sections list every section
inline; bigger splits get a single link to their SUBINDEX.md so the index
stays small. Control this via .docshelf.json:
"index_style": "auto" | "inline" | "subindex" and
"subindex_threshold_sections": 10.
If you want to keep a document whole, pass split=False.
Examples
See the examples/ directory for three concrete use cases:
examples/homelab/— original use case, hardware manuals for a home lab.examples/recipes/— a cookbook with one recipe per file.examples/research-papers/— academic PDFs with abstracts in.meta.json.
Each example shows the directory layout and the INDEX.md you'd end up with.
Optional: high-quality PDF conversion
The default engine (pymupdf4llm) is fast and good enough for ~95% of technical documents. For papers with complex tables, math, or scanned content, install the marker-pdf backend:
pip install "docshelf-mcp[high-quality]"
Then pass quality="high":
shelf.add_document("paper.pdf", category="research", title="...", quality="high")
⚠️ marker-pdf pulls in PyTorch (~2 GB) and is significantly slower (10–60 s per document on CPU). The library import is deferred — if you don't use quality="high", the dependency is never loaded.
FAQ
Why GitHub raw URLs and not embeddings / RAG? Because it's dead simple, costs nothing to host, and the AI is already good at chasing links. You can layer embedding search on top later if you want — the on-disk shape is a normal git repo.
Does this work with private repos?
Partly. The raw-URL trick needs a public repo — raw.githubusercontent.com won't serve private ones without auth. But docshelf_search and docshelf_read_document both work over MCP on private (or purely local, non-git) shelves: the model searches, then reads the exact section's content directly from the server, no raw URL required. You only lose the ability to hand a bare INDEX.md to a chat project and have it fetch by URL — with the MCP server attached, the full flow works either way. Make the doc repo public if you want the URL-fetch path too.
Do I have to use GitHub?
No. Set provider in .docshelf.json (or at init_shelf): github (default), gitlab, gitea, custom, or none. The github provider also covers GitHub Enterprise Server: a self-hosted github.<company>.com remote gets the GHES raw form (https://<host>/<owner>/<repo>/raw/<branch>/<path>) automatically. custom takes a url_template with {owner}, {repo}, {branch}, {path} placeholders, so you can point at S3, Cloudflare R2, GitLab/Gitea raw, a GHES deployment on a fully custom domain, or any static host — the generated URLs are correct everywhere, no post-processing. none renders relative links in INDEX.md, which stay navigable offline / in a local checkout.
Does it edit the source PDFs? No. P
Truncated for display — read the full file on GitHub.
Related Skills
claude-mem
96.4kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
91.6kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Understand-Anything
85.3kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.4kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
