navvi
Turn a browser task into a reusable scraper.
Install / Use
claude mcp add fellowship-dev -- npx -y github:fellowship-dev/navviIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AutomationSupported Platforms
Tags
Our assessment of navvi
navvi scores 70/100 on our quality scale, 2650th of 2,890 Automation skills we index.
Its MCP Server is 14 KB long, well organised into 13 sections with 5 code examples: a thorough specification that gives an agent plenty to work with.
It has 10 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated 9 days ago, so navvi is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 92/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
navvi compared with similar skills
All 4 of these similar skills score higher than navvi; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| navvi (this skill)by fellowship-dev | 70 | 10 | 9d ago | MCP Server |
| Agent-Reachby Panniantong | 100 | 95.3k | 2d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.9k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.3k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 86.7k | today | MCP Server |
Frequently asked questions
- How do I install navvi?
- Run
claude mcp add fellowship-dev -- npx -y github:fellowship-dev/navvi. The install tabs above show the steps for each supported agent. - Which AI agents does navvi work with?
- It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
- Is navvi safe to use?
- It is MIT-licensed and scores 92/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is navvi still maintained?
- The repository was last updated 9 days ago, so navvi is actively maintained.
Skill content
View source on GitHub- Your agent asks in plain words.
navvi "extract the title, price and availability" <url>. - Navvi compiles a scraper, fast and accurately. Code finds the candidates on the page; a model only picks among them. With Jev those picks are typed, cheap and several times faster than Haiku.
- Re-runs make zero LLM calls. The compiled scraper is a JSON file of selectors and fingerprints. Run it in cron, in CI, on Apify.
- It heals itself when the site changes. A field that moved is found again and the scraper is updated; a redesign is reported, never guessed.

Why use Jev
A scraper compiler makes a lot of small decisions: which element is the price, which control to click next, is the search done. Navvi never asks a model to write a selector or code; it enumerates the options and asks for a pick. That is exactly the shape Jev is built for.

The same 19 questions navvi asked while compiling a Hacker News search, asked again of each backend, one after another, median of three runs:
| Who decides | Time for 19 decisions | Jev is | Matched the reference | | --- | --- | --- | --- | | Jev (TypeSafe API) | 3.5 s | — | 19/19 | | Claude Haiku 4.5 over the API (AI Gateway) | 12.6 s | 3.6× faster | 19/19 | | Claude Haiku through Claude Code — navvi's default without a key | 31–38 s | 8–11× faster | 19/19 |
Same answers, a fraction of the wait. The Claude Code row is already the fast version: navvi now runs it without extended thinking, which took it from 91.6 s to 31–38 s with the same 19/19 (two sets of runs the same day; Claude Code's time varies).
Method, every batch's
latency and the captured questions:
decisions-race-provenance.json (the video, an earlier set of runs: 3.2 s against 11.0 s)
and decisions-race-claude-code.json (the
table). One task and one network location: a measurement, not a benchmark.
Turning it on is one environment variable. With TYPESAFE_API_KEY (or
AI_GATEWAY_API_KEY) set, navvi picks Jev by itself and says so:
chooser: jev (TYPESAFE_API_KEY found — fast typed decisions; text questions go to claude)
Jev picks; it does not write. The one or two free-text questions in a run (reading
your prompt, the words to type into a search box) go to a signed-in Claude Code or
Codex on your subscription, or to ANTHROPIC_API_KEY. Get a key at
typesafe.ai.
Install
Node 22+.
npm install -g navvi
npx playwright install chromium
navvi "Extract the book title, price and availability" \
https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html --browser chromium --out books.json
Run the same command again, or point it at another page of the same kind
(…/sharp-objects_997/index.html): the saved scraper answers with zero model calls.
No key and no signed-in CLI? --decider agent makes your coding agent answer the
questions itself — see Give it to your agent.
From source: git clone https://github.com/fellowship-dev/navvi.git && cd navvi && NAVVI_SKIP_BROWSER_DOWNLOAD=1 npm ci && npx playwright install chromium && npm run build, then node dist/bin/cli.js ….
What you get
-
Records on stdout or
--out data.json|data.csv, one object per item with a_sourceURL (the page it was read on) and a_startUrl(the start URL that produced it). -
A scraper you can commit:
storage/key_value_stores/scraper-cache/<key>.json. The rest ofstorage/holds browser profiles with live sessions; keep it private. -
A summary on stderr saying who answered what:
navvi: status succeeded items 2 pages 2 templates 1 cache hit no decider jev: 3 decisions, 10115 input tokens, 7.3s waiting, $0.0004 writer claude: 1 text question, 18709 input tokens, 7.7s waiting, $0.0000 healing events 0 unmapped candidates 0 unhealed 0 data: 2 records on stdout -
A run report that accounts for every page. The run summary (the
SUMMARYrecord on Apify) counts the pages that gave no row instead of healing them, each with the first 50 URLs, and a key appears only when it is non-zero:{ "status": "succeeded", "items": 918, "pages": 1000, "unhealed": 0, "blockedPages": 3, "deadPages": { "count": 41, "urls": ["https://example.com/p/retired-item"] }, "transientPages": { "count": 2, "urls": ["…"] }, "unsettledPages": { "count": 1, "urls": ["…"] }, "noPayloadPages": { "count": 4, "urls": ["…"] }, "emptyListings": { "count": 12, "urls": ["https://example.com/search?q=…"] }, "offTemplate": { "count": 5, "urls": ["https://example.com/category/…"] }, "optionalDrift": [{ "field": "listPrice", "pages": 610, "filled": 308 }] }blockedPagesis a bot challenge (the first is kept asBLOCKED_PAGE, with a screenshot);deadPagesanswered 404/410 or redirected off the template;transientPagesstill answered 5xx after a retry;unsettledPagesnever settled past a weak challenge reading;noPayloadPagesnever received the data the page loads for itself;emptyListingsare searches that found nothing;offTemplateare start URLs a pinned scraper does not match, never compiled;optionalDriftis an optional field filled on some pages and empty on others. -
Typed values when you ask:
--fields name,price:money,stock:booleanreads$ 6.990as6990andAgotadoasfalse; a value that does not coerce isnull. -
A URL list as input:
--from-url <url>fetches the pages to scrape from an endpoint (newline text or JSON), so a backend can feed the daily list. -
Optional fields (Apify actor input): a field declared
"optional": truemay be empty on a healthy page, such as a list price that exists only during a discount; replay never heals its nulls.
How it works
One compiler, cheapest evidence first. For each kind of page (a URL template):
flowchart LR
A[declared data<br/>JSON-LD, meta] --> B[the page's own<br/>JSON payloads]
B --> C[DOM candidates<br/>a model picks]
C --> D[gate + determinism]
D --> E[scraper.json]
- Tier 1 — what the page declares. JSON-LD and meta tags, read with one plain request. No browser, no model.
- Tier 2 — what the page fetches. The JSON payloads the page loads for itself, bound only when a payload is provably about this page, and never two fields to one path.
- Tier 3 — the DOM. Code enumerates candidate elements; the chooser picks one per field. Only the fields tiers 1–2 could not cover are asked.
- Close calls are questions, not guesses. When two readings compete, the
chooser is asked once, with your rules (
--rubric) in the question, and the answer is recorded with who gave it. - Replay makes no model calls. Each field carries fingerprints; when one stops matching, only that field is re-decided and the scraper is updated.
Audit a compile: --work
Add --work <dir> and the compile writes every step as a file you can read, edit
and re-run from — the spec, the sample, what each tier found, why each field was
bound (rationale.md), and a scorecard:
navvi "Extract the book title, price and availability" <url> <url> --work work/books
navvi make "<brief>" <url...> --work <dir> is the same pipeline as a resumable
driver: each stage re-runs only when the bytes it read changed, a hand-edited
artifact is never overwritten without --force, and open questions stop the run
with exit 3 until you --answer them. Both write the same scraper.json for a
record page.
Give it to your agent
Copy SKILL.md into your agent's skills (for Claude Code:
.claude/skills/navvi/SKILL.md). It tells the agent when to reach for navvi
instead of writing Playwright, the one command, how to answer questions itself
with no key (--decider agent, over stdio or with --agent-mode file and
--resume), and what each exit code means. llms.txt indexes the rest.
Choosers
Navvi asks two kinds of question. The decider answers the picks (which candidate, yes/no, a score); the writer answers free text.
| --decider | Who answers | Needs | Writes text |
| --- | --- | --- | :-: |
| jev | Jev by TypeSafe | TYPESAFE_API_KEY or AI_GATEWAY_API_KEY | no — hands text to the writer |
| claude | Claude Code on your subscription | claude installed and signed in | yes |
| codex | Codex on your subscription | codex installed and signed in | yes |
| model | Any AI SDK model | ANTHROPIC_API_KEY, or the Gateway | yes |
| agent | Your coding agent, over stdio | nothing | yes |
Default order: a Jev key → Jev; else ANTHROPIC_API_KEY → model; else a
signed-in Claude Code, then Codex; else agent. The choice and its reason print
on stderr.
Jev's writer is, in order: a signed-in claude, then codex, then a metered
key (ANTHROPIC_API_KEY, then the Gateway). A CLI that is signed out or times out
hands over to the next CLI; it never quietly starts spending on an API.
--writer pins it.
Jev's transport: with both keys set it goes over the Gateway and falls back to
the TypeSafe API if the Gateway is unavailable, saying so in the summary.
--decider-transport gateway|typesafe pins it and turns the fallback off.
--chooser <name> still works and means --decider <name>.
NAVVI_CLAUDE_MODEL (default haiku) and NAVVI_CODEX_MODEL pick the CLI model;
NAVVI_CLI_TIMEOUT_MS raises the CLI wait on a slow machine.
Exit codes
| Exit | Status | What to do |
| --- | --- | --- |
| 0 | succeeded (make: delivered) | Use the data |
| 1 | no_items_found, drift, blocked_* (make: short) | Read stderr; a redesign needs a new prompt, a login needs --profile local |
| 1 | partial | Rows were written, but a requested field was never bound and is null on every row; the status line names it (fields not found: …). Rephrase the field or pass --fields, then --force-recompile |
| 2 | configuration error | Fix the flag; the message names what is missing |
| 3 | needs_human (make: needs_answers) | Answer the parked questions and --resume, or --answer |
| 4 | budget_exhausted, model_unavailable, charge_limit | Retry later, raise the cap, or switch chooser |
Limits
- Heals a moved field or a renamed button; reports a redesign as
driftrather than guessing. - No captcha solving.
--headedhands a challenge to the person at the keyboard. - Logins use
--profile local; secrets come fromNAVVI_SECRET_<NAME>,--secrets-file, a TTY prompt or a sea
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
95.3kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.9kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
CowAgent
47.3kOpen-source personal AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Scrapling
86.7k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
