SkillAgentSearch skills...

shadow-web

Local-first Python SDK for AI agents: Shadow DOM flatten, Action Map, SchemaSnap, form fill, security scan, MCP (64–97% token reduction)

Install / Use

claude mcp add ulinycoin -- npx -y github:ulinycoin/shadow-web

If the server publishes to npm under a different name, use that package instead — check the repo README.

About this skill
🔌

MCP Server

Model Context Protocol server

Quality Score

84/100

Category

Security

Supported Platforms

Claude Code
Claude Desktop

Our assessment of shadow-web

shadow-web scores 84/100 on our quality scale, 568th of 772 Security skills we index.

Its MCP Server is 28 KB long, well organised into 85 sections with 32 code examples: a thorough specification that gives an agent plenty to work with.

It has 10 GitHub stars, so there is little community track record yet; judge it on its content.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
4/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated about 2 months ago, so shadow-web is actively maintained.
  • Our last check on 2026-09-24 found the source still online.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 97/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful.

AI review by kimi-k2.7-code on 2026-09-24. Automated pattern scan on 2026-09-24. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.

shadow-web compared with similar skills

All 4 of these similar skills score higher than shadow-web; compare them before choosing.

SkillScoreStarsUpdatedFormat
shadow-web (this skill)by ulinycoin841058d agoMCP Server
Agent-Reachby Panniantong10085.9k12d agoCLAUDE.md
headroomby headroomlabs-ai10074.0k1d agoCLAUDE.md
rufloby ruvnet10073.4ktodayCLAUDE.md
CowAgentby zhayujie10047.1ktodayCLAUDE.md

Frequently asked questions

How do I install shadow-web?
Run claude mcp add ulinycoin -- npx -y github:ulinycoin/shadow-web. The install tabs above show the steps for each supported agent.
Which AI agents does shadow-web work with?
It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
Is shadow-web safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful. It is MIT-licensed and scores 97/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is shadow-web still maintained?
The repository was last updated about 2 months ago, so shadow-web is actively maintained.
<!-- badges -->

PyPI version PyPI downloads CI Python License: MIT Stars MCP Badge

Shadow Web

Cut up to 99% of tokens from web pages before your LLM sees them — then extract structured data.
Open-source Python SDK + MCP + standalone shadow-web exec agent: flattens Shadow DOM, builds a typed Action Map, indexes rendered text on demand, heals broken selectors, and turns tables/forms/lists into clean JSON — Camoufox by default, no cloud required.

v0.4.0 — pip install "shadow-web[agent]" → shadow-web exec "…" (observe→act→verify, SchemaSnap, single-call finish).

from shadow_web.compressor import process_html

clean_html, actions, groups = process_html(raw_html)
# ✅ actions = [{"id":"1","type":"button","label":"Buy Now","group":"Checkout"}, ...]
# ✅ 164 → 46 tokens on a typical page

from shadow_web.schema_snap import parse_page

data = parse_page(clean_html)
# ✅ tables: [{columns, types, rows, total_rows}]
# ✅ forms:  [{action, method, fields: [{name, type, required, label}]}]
# ✅ lists:  [{type, items, total}]

Pain point

AI agents need to see web pages. But raw HTML is full of <script>, <style>, inline CSS, and interactive elements buried in Shadow DOM trees that Playwright can't reach. A typical Wikipedia page costs 99K tokens raw. Your LLM bill doesn't need that.

Shadow Web is what runs between the browser and the LLM: a compression layer that keeps only what matters — interactive elements, their labels, and a clean DOM skeleton. SchemaSnap then takes that clean HTML and turns it into structured data agents can actually use: table rows, form fields with validation, list items.


What you get

| Feature | Raw HTML | Playwright locators | Shadow Web | |---------|----------|-------------------|------------| | Token cost (Wikipedia) | 99,343 | — | 16,462 (−83%) | | Token cost (GitHub Trending) | 167,875 | — | 37,833 (−77%) | | Long-form via Content Index | full page | — | ~600t outline → fetch blocks | | Shadow DOM readable | ❌ | ❌ partial | ✅ flattened | | Semantic groups | ❌ | ❌ | ✅ Login / Cart / Nav | | Self-healing selectors | ❌ | ❌ | ✅ local + LLM fallback | | Consent + lazy hydration | ❌ | manual | ✅ readiness + scroll-until-content | | Catalog / feed outline ranking | ❌ | ❌ | ✅ cards / feeds before chrome | | Tables → JSON columns+rows | ❌ | ❌ | ✅ SchemaSnap | | Forms → fields with validation | ❌ | ❌ | ✅ SchemaSnap | | Lists → typed items | ❌ | ❌ | ✅ SchemaSnap | | Works offline | ✅ | ✅ | ✅ | | PyPI package | — | playwright | shadow-web |


Who this is for

| You're building … | Why Shadow Web | |-------------------|----------------| | A browser-based AI agent | Action Map + self-healing + SchemaSnap = fewer failures, structured data | | A coding agent that needs one-shot browsing | shadow-web exec — one shell call, one JSON answer (transcript stays inside) | | An MCP tool for Cursor/Claude | Built-in MCP server with 26 tools, one-command setup | | A Playwright scraper that breaks on every deploy | heal_local.py catches DOM drift without LLM cost | | A Shadow DOM-heavy app (Web components, Lit, Angular) | Read-only flatten — no React/Vue breakage | | An agent that needs data from web pages | SchemaSnap parses tables, forms, and lists into clean JSON | | Long articles, catalogs, social feeds | Rendered Text Index + ranked content_outline (cards / feeds) | | Attack surface / security recon | security_scan.py + CLI — forms, links, headers, cookies, page_class (not pentest) | | Competitor monitoring & SEO content | Shallow multi-site scans with token-bounded Action Map | | SaaS onboarding automation | AgentOps form fill — schema knows fields, LLM only picks values |


Quick install

pip install shadow-web
playwright install chromium

Extras:

pip install "shadow-web[mcp]"          # Cursor/Claude MCP server
pip install "shadow-web[server]"        # FastAPI heal API
pip install "shadow-web[agent]"         # standalone exec agent loop
pip install "shadow-web[all]"           # everything

Run the optional API locally:

export DEEPSEEK_API_KEY="..."
export SHADOW_WEB_API_KEYS="local-secret"
shadow-web-server

Remote and production requests are rejected unless SHADOW_WEB_API_KEYS is configured.

Standalone agent (shadow-web exec)

Hand a bounded browser task to the built-in observe→act→verify loop. One command in, one JSON object out — browsing transcript stays inside the sub-agent. Same thick entrypoint over MCP: agent_run(goal).

This is for answering a question / short flow — not a full-site scrape. Large volumes stay on thin tools + pagination (schema_session, content_outline(offset=…)). The envelope always reports honesty fields: complete, truncated, stop_reason, optional next_hint.

Unlike a11y-tree browser wrappers, observations are HTML Action Maps + SchemaSnap (works on any HTML, not only Playwright CDP). Default browser is Camoufox (Firefox fingerprint); Chromium is opt-in.

pip install "shadow-web[agent]"
camoufox fetch          # once — required for default --browser camoufox
# playwright install chromium   # only if you use --browser chromium
export DEEPSEEK_API_KEY="..."   # or OPENAI_API_KEY / OPENROUTER_API_KEY

shadow-web exec "What is the title of https://example.com?" \
  --model deepseek-chat
{
  "ok": true,
  "answer": "The page title is \"Example Domain\".",
  "steps": 2,
  "usage": {"prompt_tokens": 1821, "completion_tokens": 144, "total_tokens": 1965},
  "url": "https://example.com/",
  "browser": "camoufox",
  "stop_reason": "done",
  "complete": true,
  "truncated": false
}

| Flag | Default | Notes | |------|---------|--------| | --model | deepseek-chat | Any OpenAI-compatible model id | | --base-url | auto from env | DeepSeek / OpenRouter inferred from API key | | --browser | camoufox | chromium for CDP/a11y supplement path | | --max-steps | 20 | Model turns budget | | --max-tokens | 15000 | Cumulative LLM token budget (stops runaway catalog loops) | | --session | — | Persistent profile ~/.shadow-web/sessions/<name> | | --headed | off | Show the browser window | | --transcript | off | Include step list in JSON |

Dense pages (action_count > 40) auto-return minimal observations + a hint to use shadow_query / schema_page / content_outline then done — same terse pipeline as MCP, not site-specific. SPA catalogs with empty SchemaSnap should use Rendered Text Index (content_outline → content_blocks).

Programmatic: from shadow_web import run_agent_task.
Model env help: shadow-web models.


Demo

Golden path (recommended — run locally)

Full agent loop with token counts at each step:

pip install -e ".[mcp]"
playwright install chromium
python examples/golden_path/demo.py

Output: raw HTML vs navigate(minimal) + schema_session_json + shadow_query — side-by-side token table.
Playbook: examples/golden_path/CASE.md

Attack surface security scan

Automated surface mapping (not penetration testing): forms, links, page_class, HTTP security headers (HSTS, CSP, XFO, nosniff, CORS), cookie flags (Secure, HttpOnly, SameSite), Markdown/JSON reports.

pip install -e ".[mcp]"
playwright install chromium

# Single URL
python scripts/security_surface_scan.py https://example.com

# Shallow same-domain crawl + reports (--no-headers / --no-cookies to skip checks)
python scripts/security_surface_scan.py https://yoursite.com \
  --crawl-depth 1 --max-pages 20 \
  --json report.json --markdown report.md

Rule engine (importable without browser):

from shadow_web.security_scan import analyze_surface, render_markdown_report

result = analyze_surface(
    "https://app.example.com/login",
    clean_html=html,
    action_map=actions,
    page_class="Static",
)
# findings: FORM_PASSWORD_GET, HEADER_MISSING_CSP, COOKIE_MISSING_HTTPONLY, ...

Example output: examples/security_scan/localpdf-full-report.md (20 pages, 0 critical/high on public marketing layer). Header-only sample: localpdf-headers-only.json.

Does not test: XSS, SQLi, auth bypass, or deep TLS/cipher analysis. Use only on authorized targets.

AgentOps form fill

See the full Form Fill section below. Quick start:

python scripts/form_fill_demo.py https://httpbin.org/forms/post \
  --profile examples/form_fill/profile.json --json plan.json

Competitor intelligence scan

Token-bounded weekly audit for programmatic SEO / compare pages:

python scripts/localpdf_competitor_scan.py --json reports/scan.json

Playbook: examples/localpdf/CASE.md

Smoke test (install + unit tests + one live site):

bash scripts/smoke_install.sh

Compress a page (3 lines)

from shadow_web.compressor import process_html, generate_grouped_xml_map

clean_html, actions, groups = process_html(open("page.html").read())
xml_map = generate_grouped_xml_map("https://example.com", "Example", groups)
print(xml_map)

Playwright + Shadow DOM flatten

from playwright.sync_api import sync_playwright
from shadow_web.wrapper import ShadowPage

with sync_playwright() as p:
    page = p.chromium.launch(headless=True).new_page()
    page.goto("https://example.com")
    shadow = ShadowPage(page)
    _, xml_map = shadow.refresh()
    print(shadow.capture_stats)  # shadow_hosts, iframes, a11y supplement

shadow_grep — send only what the LLM needs

result = shadow.query("intent:login", fmt="terse")
# @1 button Sign in [Login Form]
# @2 input[email] Email [Login Form]

Stream a delta after clicking

shadow.refresh()               # full baseline
shadow.click("3")              # navigate
_, delta_xml = shadow.refresh(diff=True)  # only what changed

SchemaSnap — extract tables, forms, and lists

from shadow_web.schema_snap import parse_page, parse_tables

# Parse everything from a page
data = parse_page(raw_html)
# {
#   "tables": [
#     {
#       "columns": ["Product", "Price", "Stock"],
#       "types": ["string", "currency", "integer"],
#       "rows": [["Widget A", "$19.99", "150"], ...],
#       "total_rows": 12,
#       "column_count": 3
#     }
#   ],
#   "forms": [
#     {
#       "action": "/checkout",
#       "method": "POST",
#       "fields": [
#         {"tag": "input", "type": "email", "name": "email",
#          "required": true, "label": "Email Address"},
#         {"tag": "select", "name": "country", "options": [
#           {"value": "US", "label": "United States"}, ...]}
#       ]
#     }
#   ],
#   "lists": [
#     {"type": "unordered", "items": ["Apples", "Bananas"], "total": 2}
#   ]
# }

# Or export table rows directly
from shadow_web.schema_snap import export_table_json, export_table_csv

records = export_table_json(clean_html, max_rows=50)
# [{"Name": "Alice", "Age": 30}, ...]

csv_text = export_table_csv(clean_html)
# "Name,Age\nAlice,30\n..."

Sche

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars10
CategorySecurity
Updated1mo ago
Forks0

Languages

Python

Trust signals

97/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 info