shadow-web
Local-first Python SDK for AI agents: Shadow DOM flatten, Action Map, SchemaSnap, form fill, security scan, MCP (64–97% token reduction)
Install / Use
claude mcp add ulinycoin -- npx -y github:ulinycoin/shadow-webIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
SecuritySupported Platforms
Our assessment of shadow-web
shadow-web scores 84/100 on our quality scale, 568th of 772 Security skills we index.
Its MCP Server is 28 KB long, well organised into 85 sections with 32 code examples: a thorough specification that gives an agent plenty to work with.
It has 10 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated about 2 months ago, so shadow-web is actively maintained.
- Our last check on 2026-09-24 found the source still online.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 97/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful.
AI review by kimi-k2.7-code on 2026-09-24. Automated pattern scan on 2026-09-24. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
shadow-web compared with similar skills
All 4 of these similar skills score higher than shadow-web; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| shadow-web (this skill)by ulinycoin | 84 | 10 | 58d ago | MCP Server |
| Agent-Reachby Panniantong | 100 | 85.9k | 12d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.0k | 1d ago | CLAUDE.md |
| rufloby ruvnet | 100 | 73.4k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.1k | today | CLAUDE.md |
Frequently asked questions
- How do I install shadow-web?
- Run
claude mcp add ulinycoin -- npx -y github:ulinycoin/shadow-web. The install tabs above show the steps for each supported agent. - Which AI agents does shadow-web work with?
- It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
- Is shadow-web safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful. It is MIT-licensed and scores 97/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is shadow-web still maintained?
- The repository was last updated about 2 months ago, so shadow-web is actively maintained.
Skill content
View source on GitHubShadow Web
Cut up to 99% of tokens from web pages before your LLM sees them — then extract structured data.
Open-source Python SDK + MCP + standalone shadow-web exec agent: flattens Shadow DOM, builds a typed Action Map, indexes rendered text on demand, heals broken selectors, and turns tables/forms/lists into clean JSON — Camoufox by default, no cloud required.
v0.4.0 —
pip install "shadow-web[agent]"→shadow-web exec "…"(observe→act→verify, SchemaSnap, single-call finish).
from shadow_web.compressor import process_html
clean_html, actions, groups = process_html(raw_html)
# ✅ actions = [{"id":"1","type":"button","label":"Buy Now","group":"Checkout"}, ...]
# ✅ 164 → 46 tokens on a typical page
from shadow_web.schema_snap import parse_page
data = parse_page(clean_html)
# ✅ tables: [{columns, types, rows, total_rows}]
# ✅ forms: [{action, method, fields: [{name, type, required, label}]}]
# ✅ lists: [{type, items, total}]
Pain point
AI agents need to see web pages. But raw HTML is full of <script>, <style>, inline CSS, and interactive elements buried in Shadow DOM trees that Playwright can't reach. A typical Wikipedia page costs 99K tokens raw. Your LLM bill doesn't need that.
Shadow Web is what runs between the browser and the LLM: a compression layer that keeps only what matters — interactive elements, their labels, and a clean DOM skeleton. SchemaSnap then takes that clean HTML and turns it into structured data agents can actually use: table rows, form fields with validation, list items.
What you get
| Feature | Raw HTML | Playwright locators | Shadow Web |
|---------|----------|-------------------|------------|
| Token cost (Wikipedia) | 99,343 | — | 16,462 (−83%) |
| Token cost (GitHub Trending) | 167,875 | — | 37,833 (−77%) |
| Long-form via Content Index | full page | — | ~600t outline → fetch blocks |
| Shadow DOM readable | ❌ | ❌ partial | ✅ flattened |
| Semantic groups | ❌ | ❌ | ✅ Login / Cart / Nav |
| Self-healing selectors | ❌ | ❌ | ✅ local + LLM fallback |
| Consent + lazy hydration | ❌ | manual | ✅ readiness + scroll-until-content |
| Catalog / feed outline ranking | ❌ | ❌ | ✅ cards / feeds before chrome |
| Tables → JSON columns+rows | ❌ | ❌ | ✅ SchemaSnap |
| Forms → fields with validation | ❌ | ❌ | ✅ SchemaSnap |
| Lists → typed items | ❌ | ❌ | ✅ SchemaSnap |
| Works offline | ✅ | ✅ | ✅ |
| PyPI package | — | playwright | shadow-web |
Who this is for
| You're building … | Why Shadow Web |
|-------------------|----------------|
| A browser-based AI agent | Action Map + self-healing + SchemaSnap = fewer failures, structured data |
| A coding agent that needs one-shot browsing | shadow-web exec — one shell call, one JSON answer (transcript stays inside) |
| An MCP tool for Cursor/Claude | Built-in MCP server with 26 tools, one-command setup |
| A Playwright scraper that breaks on every deploy | heal_local.py catches DOM drift without LLM cost |
| A Shadow DOM-heavy app (Web components, Lit, Angular) | Read-only flatten — no React/Vue breakage |
| An agent that needs data from web pages | SchemaSnap parses tables, forms, and lists into clean JSON |
| Long articles, catalogs, social feeds | Rendered Text Index + ranked content_outline (cards / feeds) |
| Attack surface / security recon | security_scan.py + CLI — forms, links, headers, cookies, page_class (not pentest) |
| Competitor monitoring & SEO content | Shallow multi-site scans with token-bounded Action Map |
| SaaS onboarding automation | AgentOps form fill — schema knows fields, LLM only picks values |
Quick install
pip install shadow-web
playwright install chromium
Extras:
pip install "shadow-web[mcp]" # Cursor/Claude MCP server
pip install "shadow-web[server]" # FastAPI heal API
pip install "shadow-web[agent]" # standalone exec agent loop
pip install "shadow-web[all]" # everything
Run the optional API locally:
export DEEPSEEK_API_KEY="..."
export SHADOW_WEB_API_KEYS="local-secret"
shadow-web-server
Remote and production requests are rejected unless SHADOW_WEB_API_KEYS is configured.
Standalone agent (shadow-web exec)
Hand a bounded browser task to the built-in observe→act→verify loop. One command in, one JSON object out — browsing transcript stays inside the sub-agent. Same thick entrypoint over MCP: agent_run(goal).
This is for answering a question / short flow — not a full-site scrape. Large volumes stay on thin tools + pagination (schema_session, content_outline(offset=…)). The envelope always reports honesty fields: complete, truncated, stop_reason, optional next_hint.
Unlike a11y-tree browser wrappers, observations are HTML Action Maps + SchemaSnap (works on any HTML, not only Playwright CDP). Default browser is Camoufox (Firefox fingerprint); Chromium is opt-in.
pip install "shadow-web[agent]"
camoufox fetch # once — required for default --browser camoufox
# playwright install chromium # only if you use --browser chromium
export DEEPSEEK_API_KEY="..." # or OPENAI_API_KEY / OPENROUTER_API_KEY
shadow-web exec "What is the title of https://example.com?" \
--model deepseek-chat
{
"ok": true,
"answer": "The page title is \"Example Domain\".",
"steps": 2,
"usage": {"prompt_tokens": 1821, "completion_tokens": 144, "total_tokens": 1965},
"url": "https://example.com/",
"browser": "camoufox",
"stop_reason": "done",
"complete": true,
"truncated": false
}
| Flag | Default | Notes |
|------|---------|--------|
| --model | deepseek-chat | Any OpenAI-compatible model id |
| --base-url | auto from env | DeepSeek / OpenRouter inferred from API key |
| --browser | camoufox | chromium for CDP/a11y supplement path |
| --max-steps | 20 | Model turns budget |
| --max-tokens | 15000 | Cumulative LLM token budget (stops runaway catalog loops) |
| --session | — | Persistent profile ~/.shadow-web/sessions/<name> |
| --headed | off | Show the browser window |
| --transcript | off | Include step list in JSON |
Dense pages (action_count > 40) auto-return minimal observations + a hint to use shadow_query / schema_page / content_outline then done — same terse pipeline as MCP, not site-specific. SPA catalogs with empty SchemaSnap should use Rendered Text Index (content_outline → content_blocks).
Programmatic: from shadow_web import run_agent_task.
Model env help: shadow-web models.
Demo
Golden path (recommended — run locally)
Full agent loop with token counts at each step:
pip install -e ".[mcp]"
playwright install chromium
python examples/golden_path/demo.py
Output: raw HTML vs navigate(minimal) + schema_session_json + shadow_query — side-by-side token table.
Playbook: examples/golden_path/CASE.md
Attack surface security scan
Automated surface mapping (not penetration testing): forms, links, page_class, HTTP security headers (HSTS, CSP, XFO, nosniff, CORS), cookie flags (Secure, HttpOnly, SameSite), Markdown/JSON reports.
pip install -e ".[mcp]"
playwright install chromium
# Single URL
python scripts/security_surface_scan.py https://example.com
# Shallow same-domain crawl + reports (--no-headers / --no-cookies to skip checks)
python scripts/security_surface_scan.py https://yoursite.com \
--crawl-depth 1 --max-pages 20 \
--json report.json --markdown report.md
Rule engine (importable without browser):
from shadow_web.security_scan import analyze_surface, render_markdown_report
result = analyze_surface(
"https://app.example.com/login",
clean_html=html,
action_map=actions,
page_class="Static",
)
# findings: FORM_PASSWORD_GET, HEADER_MISSING_CSP, COOKIE_MISSING_HTTPONLY, ...
Example output: examples/security_scan/localpdf-full-report.md (20 pages, 0 critical/high on public marketing layer). Header-only sample: localpdf-headers-only.json.
Does not test: XSS, SQLi, auth bypass, or deep TLS/cipher analysis. Use only on authorized targets.
AgentOps form fill
See the full Form Fill section below. Quick start:
python scripts/form_fill_demo.py https://httpbin.org/forms/post \
--profile examples/form_fill/profile.json --json plan.json
Competitor intelligence scan
Token-bounded weekly audit for programmatic SEO / compare pages:
python scripts/localpdf_competitor_scan.py --json reports/scan.json
Playbook: examples/localpdf/CASE.md
Smoke test (install + unit tests + one live site):
bash scripts/smoke_install.sh
Compress a page (3 lines)
from shadow_web.compressor import process_html, generate_grouped_xml_map
clean_html, actions, groups = process_html(open("page.html").read())
xml_map = generate_grouped_xml_map("https://example.com", "Example", groups)
print(xml_map)
Playwright + Shadow DOM flatten
from playwright.sync_api import sync_playwright
from shadow_web.wrapper import ShadowPage
with sync_playwright() as p:
page = p.chromium.launch(headless=True).new_page()
page.goto("https://example.com")
shadow = ShadowPage(page)
_, xml_map = shadow.refresh()
print(shadow.capture_stats) # shadow_hosts, iframes, a11y supplement
shadow_grep — send only what the LLM needs
result = shadow.query("intent:login", fmt="terse")
# @1 button Sign in [Login Form]
# @2 input[email] Email [Login Form]
Stream a delta after clicking
shadow.refresh() # full baseline
shadow.click("3") # navigate
_, delta_xml = shadow.refresh(diff=True) # only what changed
SchemaSnap — extract tables, forms, and lists
from shadow_web.schema_snap import parse_page, parse_tables
# Parse everything from a page
data = parse_page(raw_html)
# {
# "tables": [
# {
# "columns": ["Product", "Price", "Stock"],
# "types": ["string", "currency", "integer"],
# "rows": [["Widget A", "$19.99", "150"], ...],
# "total_rows": 12,
# "column_count": 3
# }
# ],
# "forms": [
# {
# "action": "/checkout",
# "method": "POST",
# "fields": [
# {"tag": "input", "type": "email", "name": "email",
# "required": true, "label": "Email Address"},
# {"tag": "select", "name": "country", "options": [
# {"value": "US", "label": "United States"}, ...]}
# ]
# }
# ],
# "lists": [
# {"type": "unordered", "items": ["Apples", "Bananas"], "total": 2}
# ]
# }
# Or export table rows directly
from shadow_web.schema_snap import export_table_json, export_table_csv
records = export_table_json(clean_html, max_rows=50)
# [{"Name": "Alice", "Age": 30}, ...]
csv_text = export_table_csv(clean_html)
# "Name,Age\nAlice,30\n..."
Sche
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
85.9kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.0kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
ruflo
73.4k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
CowAgent
47.1kOpen-source super AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
