pulpie-mcp
MCP server that turns web pages into clean Markdown, keeping tables, code blocks, and images intact. Agents can fetch pages, save docs to your project or a global library, and crawl whole doc sites.
Install / Use
claude mcp add pinkpixel-dev -- npx -y github:pinkpixel-dev/pulpie-mcpIf the server publishes to npm under a different name, use that package instead β check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
Content & MediaSupported Platforms
Our assessment of pulpie-mcp
pulpie-mcp scores 81/100 on our quality scale, 1009th of 1,174 Content & Media skills we index.
Its MCP Server is 8.9 KB long, well organised into 18 sections with 8 code examples: a thorough specification that gives an agent plenty to work with.
It has 3 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated 2 days ago, so pulpie-mcp is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 92/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit β read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful.
AI review by kimi-k2.7-code on 2026-10-08. Automated pattern scan on 2026-10-08. It catches known dangerous patterns, not every risk β read a skill before letting an agent act on it.
pulpie-mcp compared with similar skills
All 4 of these similar skills score higher than pulpie-mcp; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| pulpie-mcp (this skill)by pinkpixel-dev | 81 | 3 | 2d ago | MCP Server |
| Agent-Reachby Panniantong | 100 | 93.2k | today | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.6k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.3k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 86.2k | today | MCP Server |
Frequently asked questions
- How do I install pulpie-mcp?
- Run
claude mcp add pinkpixel-dev -- npx -y github:pinkpixel-dev/pulpie-mcp. The install tabs above show the steps for each supported agent. - Which AI agents does pulpie-mcp work with?
- It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
- Is pulpie-mcp safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful. It is Apache-2.0-licensed and scores 92/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is pulpie-mcp still maintained?
- The repository was last updated 2 days ago, so pulpie-mcp is actively maintained.
Skill content
View source on GitHubpulpie-mcp
![]()
An MCP server that gives AI agents their own tools for turning web pages into clean Markdown. It runs the Pulpie content extraction model locally, so tables, code blocks, links, and images come through intact instead of getting summarized away.
The agent can read a page inline, save it as a .md file, crawl a whole docs section into a folder, and search everything it saved later. Over time the library turns into a local reference collection you can reuse across projects.
Why I made it
I already had a little local UI for Pulpie that I used to pull docs and articles into Markdown, and it works really well. But I was still the one doing the pulling. Most agent fetch tools run pages through a small model that summarizes or trims them, which is not what you want when the agent needs the actual API reference.
So this gives the agent the same thing I was using. It can grab full documentation whenever it needs it and save it somewhere useful, and I don't have to go do it by hand.
What it does
| Tool | What it's for |
|---|---|
| fetch_markdown | Read a page right now. The Markdown comes back inline, cut off at max_chars (40,000 by default). |
| save_markdown | Save one page as a .md file. Returns the path and a heading outline instead of the whole page, so it doesn't eat the agent's context. |
| crawl_docs | Save a documentation section as a folder of .md files plus an index.md. Up to 500 pages per crawl, 50 by default. |
| list_library | List saved docs with their source URL and fetch time, optionally filtered by a search term. |
| search_library | Search saved docs by what they say and get back the matching sections, each with its heading path, source URL, and file path. |
A few details worth knowing:
- Where files go is up to the agent. For project work it passes a folder inside the project, like
/path/to/project/DOCS/reference/uv. For general research it leavesdirectoryempty and the file goes into the global library at~/.pulpie/library/<site>/. - Every saved file has frontmatter with
title,source, andfetched, so you (and the agent) can always tell where a file came from and how old it is. - Crawling checks the sitemap first. It looks at
robots.txt, then<docs path>/sitemap.xml, then/sitemap.xml. If none of those exist, it follows links from the pages instead. It only crawls pages on the same host and under the starting path. - Search works on sections, not whole pages. Each page is split at its
#,##, and###headings, so a search foruseEffect cleanupreturns the cleanup section instead of the whole hooks page. It searches the global library by default, or any folder you pass asdirectory. - Markdown and plain-text URLs pass straight through. If a URL already serves
text/markdownortext/plain(likellms.txt), you get the file as is.
How it works
There are two pieces:
- The MCP server is what your agent launches. It's small and doesn't load the model.
- The backend is a local HTTP server that holds the Pulpie model in memory and does the actual fetching and extraction.
You never have to start the backend yourself. The first time a tool needs it, the MCP server starts it in the background and waits for the model to load. Every agent session shares that one backend, so you only ever have one copy of the model on your GPU. After 30 minutes with no requests it shuts itself down and frees the memory, and the next request starts it again.
The model is feyninc/pulpie-orange-small, a 210M parameter encoder. It uses around 420 MB of VRAM. On my RTX 4090 laptop, a cold start (including loading an already downloaded model) took about 10 seconds. After that, a single docs page saved in under a second, and an 8-page crawl of the uv docs took about 3 seconds. Most of that is network time.
Requirements
- Python 3.10 or newer
uv(recommended) orpip- An NVIDIA GPU is nice but not required. Pulpie uses CUDA if it can, then Apple MPS, then falls back to CPU.
Heads up: this pulls in PyTorch, so the install is a few GB.
I've only tested this on Linux with an NVIDIA GPU so far. It should work on macOS, and Windows is untested.
Installation
Install it as a tool so the pulpie-mcp command is on your PATH:
uv tool install git+https://github.com/pinkpixel-dev/pulpie-mcp
Or from a local clone:
git clone https://github.com/pinkpixel-dev/pulpie-mcp
cd pulpie-mcp
uv tool install .
The model downloads from Hugging Face the first time the backend starts.
Adding it to your agent
Claude Code
claude mcp add --scope user pulpie -- pulpie-mcp
Codex
In ~/.codex/config.toml:
[mcp_servers.pulpie]
command = "pulpie-mcp"
Other MCP clients
Most clients take a JSON config like this:
{
"mcpServers": {
"pulpie": {
"command": "pulpie-mcp"
}
}
}
If you add environment variables (see below), put them in the client's env block for this server.
Using it
You don't need special prompts. Just ask for what you want:
- "Pull the uv docs into this project's
DOCS/referencefolder." - "Read the asyncio queue docs and tell me how
join()works." - "Do we already have the FastAPI docs saved somewhere?"
- "Search the library for how FastAPI handles background tasks."
The server tells the agent to save project docs inside the project and general research in the global library, so it usually picks the right spot on its own.
Configuration
Everything is optional.
| Variable | Default | What it does |
|---|---|---|
| PULPIE_URL | http://127.0.0.1:8787 | Where the backend runs. Only local URLs get auto-started. |
| PULPIE_HOME | ~/.pulpie | Holds the global library, the search index (search.db), and backend.log. |
| PULPIE_LIBRARY | $PULPIE_HOME/library | Where saved docs go when no directory is given. |
| PULPIE_ALLOW_PRIVATE | off | Set to 1 to allow localhost and LAN addresses. |
| PULPIE_IDLE_TIMEOUT | 1800 | Seconds without requests before the backend exits. 0 keeps it running. |
| PULPIE_START_TIMEOUT | 600 | How long to wait for the backend to start. The first run includes the model download. |
Private addresses
By default the server won't fetch localhost, 127.0.0.1, LAN addresses like 192.168.x.x, or cloud metadata addresses. Agents follow links they find on the open web, so I'd rather block those than have a random page send the agent poking around your network. This is checked on every redirect too.
If you want to pull docs from a local dev server, set PULPIE_ALLOW_PRIVATE=1.
Commands
pulpie-mcp # run the MCP server over stdio (your agent does this)
pulpie-mcp serve # run the backend in the foreground, handy for debugging
pulpie-mcp status # show whether the backend is running, and on which device
pulpie-mcp stop # stop the backend
If something goes wrong with the backend, check ~/.pulpie/backend.log.
Limitations
- Pages that render entirely in JavaScript come back mostly empty, because the backend fetches the HTML the server sends and doesn't run a browser. Most docs sites (Docusaurus, MkDocs, VitePress, Sphinx, Next.js) render on the server, so this hasn't been an issue for docs so far.
- Code blocks are indented, not fenced. Pulpie converts HTML to Markdown with
html2text, which writes code as 4-space indented blocks without a language tag. It's valid Markdown, just not as pretty. - Search is keyword search, not semantic search. It uses SQLite full-text search with stemming, so
cachingfindscacheandreadOnlyHintfindsread_only_hint. A question phrased in totally different words than the docs use can still miss. If nothing matches every word, the results say they only match some of the words. - No PDFs. Non-HTML pages other than Markdown and plain text are rejected.
- Saved pages are untrusted web content. Anything an agent reads from them is data, not instructions.
Model license
[!IMPORTANT] This project's code is Apache 2.0, but the Pulpie model it downloads,
feyninc/pulpie-orange-small, is licensed CC BY-NC 4.0. That means non-commercial use only. ThepulpiePython library itself is Apache 2.0. If you want to use this in a commercial product, check the model license with Feyn first.
pulpie-mcp doesn't bundle or redistribute the model weights. They're downloaded from Hugging Face on your machine.
Development
uv sync
uv run pytest
The integration test runs the real backend and model against a live page, so it's skipped unless you opt in:
PULPIE_INTEGRATION=1 uv run pytest tests/test_integration.py
License
Apache 2.0. See LICENSE.
Pulpie is made by Feyn.
Made with π by Pink Pixel
Related Skills
Agent-Reach
93.2kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu β one CLI, zero API fees.
headroom
74.6kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
CowAgent
47.3kOpen-source personal AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Scrapling
86.2kπ·οΈ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit β see the Safety scan above for what the skill file itself contains.