stalin
☭ Turn any website into a self-healing, typed JSON API you can query — with filters, or just ask in plain English. Local-first, MCP-native.
Install / Use
claude mcp add DreadpiratePickles -- npx -y github:DreadpiratePickles/stalinIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AutomationSupported Platforms
Our assessment of stalin
stalin scores 75/100 on our quality scale, 2603rd of 2,869 Automation skills we index.
Its MCP Server is 22 KB long, well organised into 25 sections with 19 code examples: a thorough specification that gives an agent plenty to work with.
It has 3 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated 20 days ago, so stalin is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 92/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
stalin compared with similar skills
All 4 of these similar skills score higher than stalin; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| stalin (this skill)by DreadpiratePickles | 75 | 3 | 20d ago | MCP Server |
| Agent-Reachby Panniantong | 100 | 91.8k | 20d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.5k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.2k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.9k | 1d ago | MCP Server |
Frequently asked questions
- How do I install stalin?
- Run
claude mcp add DreadpiratePickles -- npx -y github:DreadpiratePickles/stalin. The install tabs above show the steps for each supported agent. - Which AI agents does stalin work with?
- It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
- Is stalin safe to use?
- It is MIT-licensed and scores 92/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is stalin still maintained?
- The repository was last updated 20 days ago, so stalin is actively maintained.
Skill content
View source on GitHub☭ stalin
Turn any website into a self-healing, typed JSON API.
He seizes the means of extraction.
</div>Every scraper you have ever written is already dead. It just doesn't know it yet.
Some Tuesday, a frontend dev renames .score to .karma, your pipeline fills
with nulls, and you find out three days later when a dashboard flatlines.
stalin does not tolerate this. You describe what you want in plain English, once. stalin compiles that into CSS selectors, watches them like a paranoid quartermaster, and when the website changes its layout — and it will — stalin detects the drift in microseconds, re-locates your data, proves the fix against five deterministic gates, and keeps shipping typed JSON like nothing happened. The selector was purged. The schema survives. The API never changes shape without your explicit order.
Your scraper doesn't break. It gets reeducated.
The 60-second demo
$ pip install stalin-scraper
$ stalin init
$ stalin add https://news.ycombinator.com -n stories \
--item "each story row on the front page" \
-f "title: str the headline text of each story
url: url the link the headline points to
points: int | null the upvote count
comments: int | null number of comments"
✓ fetched news.ycombinator.com 200 · 41 KB · 0.3s
✓ title .titleline > a verified 30/30 via heuristic · 2 fallbacks
✓ url .titleline a @href verified 30/30 via heuristic · 2 fallbacks
✓ points .score verified 29/30 via heuristic · 2 fallbacks
✓ comments .subline a:last-child verified 30/30 via llm · 2 fallbacks
✓ wrote sources/stories.yml, .stalin/lock.json (schema v1.0.0, 8 fallbacks)
$ stalin run stories | jq '.items[0]'
{ "title": "Show HN: …", "url": "https://…", "points": 312, "comments": 148 }
$ stalin serve
stalin API → http://127.0.0.1:8411
GET /v1/stories typed JSON, schema v1.0.0
GET /openapi.json OpenAPI 3.1
GET /healthz per-source status
And then — weeks later, when the site ships a redesign — the moment this tool
exists for. This is a real transcript (demo site renamed every class and
swapped <h2> for <div>):
$ stalin run stories
⚠ drift detected · stories.points
selector .score
signals zero-match
✓ heal title .story .headline ──→ .story .hdr via heuristic
✓ heal url .headline a @href ──→ .hdr a @href via heuristic
✓ heal points .score ──→ .karma via heuristic
✓ heal comments .comments ──→ .replies via heuristic
✓ stories 20 items schema v1.0.0 (4 fields healed)
$ stalin history stories
11:16:16Z ✓ heal points .score ──→ .karma via heuristic · zero-match
11:16:28Z ✓ confirmed points
Four fields, one redesign, zero human intervention, zero schema changes,
milliseconds of healing. The comrades downstream consuming /v1/stories
never noticed anything.
Why this exists
| | |
|---|---|
| 🩹 It never breaks silently | Five drift signals run on every scrape — including the nasty case where your selector still matches something, just the wrong something. Detection costs microseconds, not LLM calls. |
| 📐 The schema is law | Fields are typed (int, url, datetime, list[str], enum[…]), validated through pydantic on every run, and versioned with semver. Healing may change where data comes from — never what shape it has. Breaking changes require an explicit stalin schema bump --major. Yes, it makes you type --major. That's the point. |
| 🏠 Local-first, zero API keys | The healer runs on your machine through Ollama. No cloud, no per-token bill, no data leaving the building. |
| 🧠 The LLM is optional | Compile and healing run a ladder: stored fallbacks → DOM-statistics heuristics → local LLM. On most pages the heuristics do everything in milliseconds and the model is never consulted. No Ollama at all? You still get full detection + alerting. |
Any website → a live, queryable API ⚡ new in 0.2
A static_feed gives you one fixed page as JSON. But most of the web you'd
actually want as an API is parameterized — a search box, a profile page,
a lookup by ID. stalin turns those into live typed endpoints. Put a {param}
in the URL and give one example to compile against:
$ stalin add "https://quotes.toscrape.com/tag/{tag}/" -n quotes \
--item "each quote block on the page" \
--example tag=love \
-f "text: str the quote text itself
author: str who said it
tags: list[str] the topic tags on the quote"
✓ fetched quotes.toscrape.com/tag/love/ 200 · 12 KB
✓ text .quote .text verified 10/10 via heuristic
✓ author .author verified 10/10 via heuristic
✓ tags .tags .tag verified 10/10 via heuristic
live lookup — params: tag
That website has no API. It does now:
$ stalin run quotes --param tag=courage --json | jq '.count'
2
$ stalin serve
GET /v1/quotes?tag=humor # live, typed, on demand
GET /openapi.json # the param is a documented query parameter
$ curl 'http://127.0.0.1:8411/v1/quotes?tag=love' | jq '.items[0].author'
"André Gide"
$ curl 'http://127.0.0.1:8411/v1/quotes' # missing required param
{"error": "missing required param(s): tag", "params": {"tag": "str"}}
And it becomes a typed tool for your agents. Every parameterized source is
auto-exposed over MCP as lookup_<name>(param=…) with a generated input
schema — so Claude (or any agent) calls lookup_quotes(tag="stoicism") and
gets back schema-guaranteed JSON, no scraping code in the agent, no HTML in the
prompt. That's stalin's answer to "why not just let the agent scrape?" — the
tool's contract stays honest even as the site churns.
Lookups self-heal too — the hard version
A live lookup returns different content for every query, so "zero results" can
mean the query was empty or the site broke. stalin distinguishes them: you
declare a not_found signal for legitimately-empty pages (never healed
against, same rule as block pages), and healing is anchored to a heal
fixture — the known-good example params you compiled with. When a real query
comes back unexpectedly empty, stalin re-verifies selectors against the fixture
(stable, known structure), heals there, and re-applies the fix to your query.
Per-query variation never gets mistaken for drift.
Proven live: pointed at a page, compiled, then renamed every CSS class and
swapped the tags — the very next ?param=… request healed three fields against
the fixture and returned correct typed data, in one round, no human touch.
What works today, honestly
live_lookup runs on static-HTML pages right now. Path params (/{id}/)
and query params (?q=…) both work. What it does not do yet, and won't
pretend to: JavaScript-rendered pages (the Playwright engine seam exists but
isn't built), pagination/infinite-scroll, and anything behind a login or
CAPTCHA — those stay refused, by design. It does public, unauthenticated,
static surfaces. That covers a huge amount of "this site should've had an
API" — and none of the stuff that gets you sued.
How healing works
drift detected on field F
├─ Rung 1: stored fallbacks compile-time alternates, ~ms, no LLM
├─ Rung 2: heuristic re-location DOM statistics scored against the
│ field's fingerprint, ~ms, no LLM
├─ Rung 3: LLM re-location local model, sentinel text protocol —
│ works with any Ollama model
└─ Rung 4: broken serve last-good data, exit 3, tell you
The proposers differ per rung. The judge never changes: every candidate selector faces five deterministic acceptance gates —
- Schema — every extracted value must cast to the declared type
- Cardinality — match counts must stay near the historical baseline
- Shape — values must match the field's learned value-pattern (or migrate to a consistent new one, which is recorded)
- Anchor — if an old known-good value still exists on the page, the new selector must capture it exactly (substring lookalikes are rejected)
- Disjointness & robustness — no annexing another field's selector, no
position-brittle
:nth-child(7)nonsense, no auto-generated class hashes
The LLM proposes. The gates dispose. No model opinion is ever trusted about
its own output — every accepted heal was executed against the real DOM and
survived all five gates. Accepted heals are applied immediately (data keeps
flowing) but marked healed-unconfirmed until two clean runs promote them;
a re-drift inside that window reverts the heal and flags the field instead of
thrashing on A/B-tested sites.
Everything is recorded. git diff .stalin/lock.json shows every selector the
healer has ever touched, and .stalin/history/*.jsonl is an append-only,
line-per-event audit log. Rewriting history is for websites, not for your
data pipeline.
Install
pip install stalin-scraper
Optional but recommended — a local model for the healing rung:
# any ollama model works; small ones are fine (the gates do the hard part)
ollama pull qwen3:1.7b
Then check your environment:
stalin doctor
Usage
| Command | What it does |
|---|---|
| stalin init | Scaffold a project (stalin.yml, sources/, .stalin/) |
| stalin add <url> -n NAME -f "…" | Compile a new source: fetch → generate selectors → verify → save |
| stalin add "<url/{param}>" --example param=v … | Compile a live_lookup: any templated URL → a queryable API |
| stalin run SOURCE --param k=v | Run a live lookup for specific params |
| stalin run SOURCE -w "f__op=v" --sort -f --fields a,b --limit N | Filter/sort/select/paginate results |
| GET /v1/SOURCE?f__gt=1&sort=-f&fields=a,b&limit=N&q=… | Query over HTTP (see Query it) |
| stalin ask SOURCE "plain-English question" | Let the local LLM compile a query for you |
| stalin run [SOURCE…] | Extract now. Auto-heals on drift. JSON to stdout when piped |
| stalin run --no-heal | Detect drift, report, exit 3 — never heal (CI mode) |
| stalin heal [SOURCE[.FIELD]] | Force a heal pass (includes the LLM rung) |
| stalin serve [--refresh 15m] | HTTP API over cached snapshots + OpenAPI 3.1 |
| stalin watch | Foreground scheduler: run each source on its schedule: |
| stalin mcp | MCP server on stdio — plug your sources into Claude/any agent |
| stalin schema show/bump | Print the JSON Schema; explicitly version-bump the contract |
| stalin history SOURCE[.FIELD] | The heal/drift/confirm timeline |
| stalin snapshot SOURCE | Refresh the reference HTML snapshot |
| stalin doctor | Environment + per-source health check |
Exit codes are a contract (cron/CI friendly): 0 ok · 1 config error ·
2 fetch blocked · 3 drift unhealed (stale data served) · 4 schema
contract breach.
Declaring a source
stalin add writes this file — or write it yourself and let stalin compile it:
# sources/stories.yml — the INTENT. Hand-editable, never touched by the healer.
name: stories
url: https://news.ycombinator.com
schedule: 15m
item: each story row on the front page # natural language!
fields:
title:
type: str
desc: the headline text of each story
points:
type: int | null
desc: the upvote count; job postings have none
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
91.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.5kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
CowAgent
47.2kOpen-source personal AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Scrapling
85.9k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
