SkillAgentSearch skills...

trawl

Selective web content extraction for AI agents — URL + query returns only the chunks that matter (Python library + MCP server)

Install / Use

claude mcp add bbulb -- npx -y github:bbulb/trawl

If the server publishes to npm under a different name, use that package instead — check the repo README.

About this skill
🔌

MCP Server

Model Context Protocol server

Quality Score

78/100

Supported Platforms

Claude Code
Claude Desktop

Our assessment of trawl

trawl scores 78/100 on our quality scale, 802nd of 966 AI & Machine Learning skills we index.

Its MCP Server is 22 KB long, well organised into 30 sections with 15 code examples: a thorough specification that gives an agent plenty to work with.

It has 3 GitHub stars, so there is little community track record yet; judge it on its content.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
3/20
Freshness
11/15

Maintenance, license and trust

  • The repository was last updated about 3 months ago. That is recent enough to be usable, but agent tooling moves fast, so check the instructions against your agent's current version.
  • Our last check on 2026-09-12 found the source still online.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 90/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

trawl compared with similar skills

All 4 of these similar skills score higher than trawl; compare them before choosing.

SkillScoreStarsUpdatedFormat
trawl (this skill)by bbulb7833mo agoMCP Server
claude-memby thedotmack10099.5ktodayCLAUDE.md
Agent-Reachby Panniantong10095.7k3d agoCLAUDE.md
Understand-Anythingby Egonex-AI10085.9k1d agoCLAUDE.md
headroomby headroomlabs-ai10075.0ktodayCLAUDE.md

Frequently asked questions

How do I install trawl?
Run claude mcp add bbulb -- npx -y github:bbulb/trawl. The install tabs above show the steps for each supported agent.
Which AI agents does trawl work with?
It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
Is trawl safe to use?
It is MIT-licensed and scores 90/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is trawl still maintained?
The repository was last updated about 3 months ago. That is recent enough to be usable, but agent tooling moves fast, so check the instructions against your agent's current version.

trawl

CI Python 3.10+ License: MIT

Selective web content extraction for AI agents. Give trawl a URL and a natural-language query; it fetches the page, extracts the main content, chunks it, embeds the chunks with a local bge-m3 model, and returns only the handful most relevant to the query.

The point is to let an agent "read a web page" by reading only the ~1,000 tokens that matter, instead of dumping 50k+ tokens of page content into its context.

from trawl import fetch_relevant

r = fetch_relevant("https://en.wikipedia.org/wiki/Yi_Sun-sin",
                   "who did Yi Sun-sin defeat at Myeongnyang")
for c in r.chunks:
    print(f"[{c['score']:.2f}] {c['heading']}\n    {c['text'][:120]}")

Why trawl?

Most "read this page" tools fall into two camps:

  1. Full-page dumpers (Jina Reader, Firecrawl markdown) — faithful but dump the entire page into your context window. A 50k-token documentation page becomes 50k tokens of input regardless of what you actually wanted to know.
  2. LLM-driven extractors (Firecrawl /extract) — ask an LLM to pull structured fields, which needs a strong model, is slow, and still ships the full page to the model internally.

trawl takes a different angle: query-aware dense retrieval over the extracted markdown. The heavy lifting is a small, fast local embedding model (bge-m3), not an LLM. You get back the 5-12 chunks that matter for your query, at ~1k tokens of output.

Benchmark vs Jina Reader (12 cases)

| Mode | Avg tokens returned | vs Jina | Ground-truth pass | |---|---|---|---| | trawl-base | 1,177 | 23× fewer | 11/12 | | trawl-cached (with profile) | 1,004 | 30× fewer | 10/11 | | Jina Reader | 27,506 | (baseline) | 12/12 |

trawl wins on every token-efficiency axis and runs entirely on your own infrastructure. In exchange you pay a real cost elsewhere:

External: WCXB dev (1,497 pages)

Beyond the internal 15-case parity matrix, trawl's extraction stage is cross-validated against the WCXB public benchmark (CC-BY-4.0, 1,497 dev pages across 7 page types).

| Extractor | F1 | |-----------------------------------|--------| | trawl (html_to_markdown) | 0.818 | | Trafilatura (same environment) | 0.750 |

(0.818 as of v0.4.6 — rs-trafilatura candidate default-on plus selector-scoring fixes; v0.4.5 measured 0.777.)

Per-page-type breakdown and error counts: see benchmarks/wcxb/README.md and run the benchmark locally to regenerate.

When not to use trawl

  • You want the whole page verbatim. Selective retrieval is the point; if your downstream task needs faithful full-page markdown (archival, translation, full-text search indexing), Jina Reader or Firecrawl's markdown mode is the right tool.
  • Low-friction setup matters more than token efficiency. Jina is curl https://r.jina.ai/<url> — one HTTP call, no local state. trawl needs a Python environment, Chromium via Playwright, and a running bge-m3 embedding server you host yourself.
  • Latency-sensitive first-visit calls. Jina's CDN ~3s vs trawl's ~9s on the first fetch (Playwright + stealth + embedding). With a cached profile trawl's subsequent fetches to the same host drop, but the first visit is always slower.
  • Sites behind active anti-bot (Cloudflare Turnstile with proof-of-work, DataDome). trawl's local playwright-stealth defeats passive JS challenges only; commercial services that pay for anti-bot infrastructure will get those pages where trawl can't.
  • No query, just "read this". trawl requires a query to rank against (unless a cached profile exists). For "summarise whatever this page is about", a full-page dumper is a better fit.

What's in the box

  • Adaptive fetcher routing — API-first fetchers for YouTube, Wikipedia, Stack Exchange, GitHub, and arXiv PDFs; Playwright + playwright-stealth fallback for everything else.
  • Three-way extraction — Trafilatura (precise + recall) and BeautifulSoup heuristics race; the longest result wins. This covers articles, pricing pages, and lists without per-site rules.
  • Heading-aware chunker — preserves heading context on every chunk and keeps tables intact. Falls back to sentence-level chunking for PDF-style single-blob inputs.
  • Repeating-record chunking — when the rendered DOM contains a run of sibling elements with the same structural signature (job listings, news cards, product rows), each record becomes its own atomic chunk so retrieval ranks them individually instead of fragmenting mid-record.
  • Raw passthrough for JSON / XML / RSS / Atom — URLs with those suffixes (or endpoints that answer Content-Type: application/json on a HEAD probe) are returned byte-for-byte up to TRAWL_PASSTHROUGH_MAX_BYTES (default 256 KB). No embedding, no query required.
  • bge-m3 dense retrieval with an OpenAI-compatible embedding endpoint. Adaptive top-k based on page size. If the embedding endpoint is unavailable, trawl falls back to BM25 lexical ranking and returns a warning in the result payload instead of failing the entire fetch. This degraded mode is meant for operational continuity; quality is best with the bge-m3 embedding service running.
  • Cross-encoder reranking (bge-reranker-v2-m3) on the top 2× candidates. Falls back gracefully to cosine-only if the reranker server is down.
  • Chunk budget for longform pages (default on). When a page produces more chunks than TRAWL_CHUNK_BUDGET (default 100), a BM25 prefilter keeps the top-N and drops the rest before embedding. Cuts retrieval cost on Wikipedia / arXiv / long manpage scale pages (~69% retrieval_ms.p95 reduction on longform fixtures; rank-1 identity preserved). Opt out via TRAWL_CHUNK_BUDGET=0.
  • Optional HyDE query expansion for queries where the literal words don't match the page vocabulary. Off by default.
  • VLM page profiling (optional) — when the same site is visited repeatedly, trawl can ask a vision LLM to propose a CSS selector that scopes future fetches to the article region. Cached per host.
  • Indirect prompt-injection defense (default on) — fetched content is scanned (model-free): invisible Unicode tag characters are stripped, and instruction-like text — whether in visible chunks or in CSS-hidden / off-screen nodes — is flagged with suspicious_injection / suspicious_hidden and a scan warning, so the consuming agent gets an untrusted-content signal instead of silently relayed instructions. Disable via TRAWL_INJECTION_SCAN=0.
  • stdio MCP server exposing fetch_page and profile_page tools for Claude Code, Claude Desktop, and any MCP-compatible client.

Project layout

src/trawl/                  pipeline library
  pipeline.py               fetch_relevant() entry point
  chunking.py               heading + table preserving chunker
  records.py                repeating-sibling record detection + sentinels
  retrieval.py              bge-m3 cosine retrieval, adaptive k
  reranking.py              bge-reranker-v2-m3 cross-encoder
  extraction.py             Trafilatura + BeautifulSoup three-way
  hyde.py                   optional query expansion
  telemetry.py              opt-in JSONL telemetry
  profiles/                 VLM-based page profiling (optional)
  fetchers/                 per-site API-first adapters
    playwright.py, pdf.py, passthrough.py, youtube.py,
    wikipedia.py, github.py, stackexchange.py

src/trawl_mcp/              MCP server (stdio default, --http opt-in)
tests/                      unit tests + 15-case parity matrix
benchmarks/                 trawl vs Jina, VLM profile eval
examples/                   MCP client config snippets

See ARCHITECTURE.md for the design rationale behind every component, per-case performance, and known limitations.

Requirements

  • Python 3.10+
  • Chromium (installed via Playwright)
  • A running bge-m3 embedding server with an OpenAI-compatible /v1/embeddings endpoint. The reference setup is llama-server loaded with a bge-m3 GGUF, listening on http://localhost:8081. Any OpenAI-compatible embedding endpoint works if you override TRAWL_EMBED_URL.

Optional:

  • bge-reranker-v2-m3 on :8083 for cross-encoder reranking (graceful fallback if absent)
  • A small utility LLM on :8082 for HyDE (off by default)
  • A vision LLM on :8080 for profile_page (only needed if you use the profiling feature)

Reference llama-server commands

The reference local setup runs four llama-server processes. These are the canonical flag sets validated against the parity matrix and the coding agent_patterns shard. Adjust -ngl, --ctx-size, and --parallel to your hardware.

# :8081 — bge-m3 embeddings (REQUIRED for retrieval)
llama-server -m ~/models/bge-m3-Q8_0.gguf \
  --embeddings --pooling cls \
  --port 8081 -ngl 99 --ctx-size 8192 -ub 2048 -b 2048

# :8083 — bge-reranker-v2-m3 cross-encoder (optional, graceful fallback)
llama-server -m ~/models/bge-reranker-v2-m3-Q8_0.gguf \
  --reranking --pooling rank \
  --port 8083 -ngl 99 --ctx-size 65536 \
  --parallel 4 -ub 2048 -b 2048

# :8082 — small utility LLM for HyDE (optional, off by default)
llama-server -m ~/models/gemma-3-4b-it-Q4_K_M.gguf \
  --port 8082 -ngl 99 --ctx-size 4096

# :8080 — vision LLM for profile_page (optional, manual-trigger only)
llama-server -m ~/models/<vision-model>.gguf --mmproj ~/models/<mmproj>.gguf \
  --port 8080 -ngl 99 --ctx-size 8192

Install

The reference setup uses a dedicated conda/mamba environment (environment.yml creates it):

mamba env create -f environment.yml    # creates `trawl` env with deps
mamba run -n trawl playwright install chromium

Copy .env.example → .env if you need to override any default endpoints; every variable is optional.

All commands below assume the trawl mamba environment: either activate it with mamba activate trawl or prefix commands with mamba run -n trawl.

Runtime health check

Use trawl-doctor to check the local runtime before wiring trawl into an MCP client:

trawl-doctor
# or
python -m trawl.diagnostics --json

The doctor checks Python, Playwright Chromium, cache-path writability, the embedding endpoint, the optional reranker endpoint, and optional VLM profile configuration. Embedding is required for dense retrieval; reranker and VLM are optional.

Usage

As a Python library

from trawl import fetch_relevant

result = fetch_relevant(
    "https://ko.wikipedia.org/wiki/이순신",
    "이순신 직업 생년월일 주요 업적",
)

print(f"fetcher={result.fetcher_used}  latency={result.total_ms}ms")
print(f"compression={result.compression_ratio}x")
for chunk in result.chunks:
    print(f"[{chunk['score']:.3f}] {chunk['heading']}")
    print(f"    {chunk['text'][:200]}")

fetch_relevant never raises. On failure it returns a PipelineResult with an empty chunks list and a non-empty error — check result.error before consuming result.chunks.

As an MCP server (stdio)

python -m trawl_mcp
# or, if the console script is on PATH:
trawl-mcp

The server exposes two tools:

fetch_page — retrieval.

| Field | Type | Required | Default | Description | |---|---|---|---|---| | url | string | yes | — | Target URL. .pdf URLs or URLs containing /pdf/ route through the PDF path | | query | string | no | — | The user's question/topic. Required when no cached profile

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars3
CategoryAI
Updated3mo ago
Forks0

Languages

Python

Trust signals

90/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 low1 info