crawlemoon
Advanced AI-native web crawling platform for the agentic era. Exposes 55 production-grade Model Context Protocol (MCP) tools for deep analysis, automated REST/GraphQL discovery, browser session recording, anti-detection stealth, and smart extraction using any LLM.
Install / Use
claude mcp add razavioo -- npx -y github:razavioo/crawlemoonIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AutomationSupported Platforms
Our assessment of crawlemoon
crawlemoon scores 81/100 on our quality scale, 1582nd of 2,166 Automation skills we index.
Its MCP Server is 8.4 KB long, well organised into 14 sections with 7 code examples: a thorough specification that gives an agent plenty to work with.
It has 3 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated 27 days ago, so crawlemoon is actively maintained.
- Our last check on 2026-09-20 found the source still online.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 92/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
crawlemoon compared with similar skills
All 4 of these similar skills score higher than crawlemoon; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| crawlemoon (this skill)by razavioo | 81 | 3 | 27d ago | MCP Server |
| Agent-Reachby Panniantong | 100 | 86.1k | 13d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.1k | today | CLAUDE.md |
| rufloby ruvnet | 100 | 73.5k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.2k | today | CLAUDE.md |
Frequently asked questions
- How do I install crawlemoon?
- Run
claude mcp add razavioo -- npx -y github:razavioo/crawlemoon. The install tabs above show the steps for each supported agent. - Which AI agents does crawlemoon work with?
- It is written for Claude Code, Claude Desktop, Cursor and OpenAI Codex, as a MCP Server file. Other agents that read the same format can often use it too.
- Is crawlemoon safe to use?
- It is MIT-licensed and scores 92/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is crawlemoon still maintained?
- The repository was last updated 27 days ago, so crawlemoon is actively maintained.
Skill content
View source on GitHubCrawlemoon MCP Server
<p align="center"> <img src="https://raw.githubusercontent.com/razavioo/crawlemoon/main/assets/hero.png" alt="Crawlemoon MCP Server — free, AI-native web crawling for the agent era" width="100%"/> </p> <p align="left"> <img alt="python 3.10+ · pypi 1.1.0 · MIT · MCP-native · code style black" src="https://raw.githubusercontent.com/razavioo/crawlemoon/main/assets/badges.png" height="22"/> </p>A free, open-source MCP server that gives any agent (Claude Code, Cursor, Windsurf, …) 55 production-grade tools for the full web-crawling stack: deep analysis, stealth, API discovery, session recording → runnable crawler, smart extraction. No proprietary API. No per-request fee.
<p align="center"> <img src="https://raw.githubusercontent.com/razavioo/crawlemoon/main/assets/features.png" alt="Crawlemoon capabilities — deep analysis, stealth, record→crawler, smart extraction" width="100%"/> </p>Quick start
<p align="center"> <img src="https://raw.githubusercontent.com/razavioo/crawlemoon/main/assets/install.png" alt="Three install paths — uvx, pipx, pip" width="100%"/> </p>The recommended path needs no install — uvx runs straight from PyPI:
{
"mcpServers": {
"crawlemoon": {
"command": "uvx",
"args": ["crawlemoon"]
}
}
}
Requires
uv. Install once:curl -LsSf https://astral.sh/uv/install.sh | sh. Or usepipx run crawlemoon/pip install crawlemooninstead.
Where to put that JSON: Cursor → Settings → MCP. Claude Code → ~/.config/claude/mcp_settings.json. Windsurf → Settings → MCP Servers.
How it works
<p align="center"> <img src="https://raw.githubusercontent.com/razavioo/crawlemoon/main/assets/architecture.png" alt="Agent → Crawlemoon → Browser/HTTP/Proxy → target web" width="100%"/> </p>Your agent talks to Crawlemoon over the Model Context Protocol. Crawlemoon owns a hardened browser pool, an HTTP stack with TLS fingerprinting, and a rotating proxy pool. While it fetches pages, it captures network traffic, reads scripts, and introspects schemas — so the agent gets clean structured data, not raw HTML.
What's in the box
A short list — see the source for the full set of 55 tools.
| Group | Tools |
|---|---|
| Deep analysis | deep_analyze, discover_apis, introspect_graphql, analyze_websocket, analyze_auth, detect_protection, detect_technology |
| Stealth | stealth_request, configure_proxies, configure_rate_limit, add_proxy, test_proxy |
| Record → crawler | record_session, stop_recording, export_recording, generate_crawler |
| Extraction | smart_extract, extract_article, extract_tables, extract_links, extract_forms, extract_metadata, convert_to_markdown |
| Page interaction | take_screenshot, fill_form, wait_and_extract, compare_pages, measure_performance, check_accessibility, get_dom_tree |
| Sessions & cache | save_session, load_session, get_cookies, get_storage, clear_cache, get_cache_stats |
| Advanced (opt-in) | execute_js, execute_cdp, deobfuscate_js, extract_from_js, solve_captcha |
Smart extraction — bring any LLM, including free ones
smart_extract works without any API key using pattern matching. Plug in any OpenAI-compatible endpoint for higher accuracy — including FREE tiers:
# OpenRouter (free models exist)
CRAWLEMOON_LLM_PROVIDER=openrouter
CRAWLEMOON_LLM_API_KEY=sk-or-v1-xxx
CRAWLEMOON_LLM_MODEL=meta-llama/llama-3.2-3b-instruct:free
# Groq (free, very fast)
CRAWLEMOON_LLM_PROVIDER=groq
CRAWLEMOON_LLM_API_KEY=gsk_xxx
# Local Ollama (no key needed)
CRAWLEMOON_LLM_PROVIDER=ollama
CRAWLEMOON_LLM_MODEL=llama3.2
Together, DeepSeek, Mistral, Fireworks, and standard OpenAI also work via CRAWLEMOON_LLM_BASE_URL.
Configuration
| Variable | Default | Notes |
|---|---|---|
| CRAWLEMOON_HEADLESS | true | Run browser without UI |
| CRAWLEMOON_BROWSER | chromium | chromium / firefox / webkit |
| CRAWLEMOON_POOL_SIZE | 5 | Max concurrent browsers |
| CRAWLEMOON_NAV_TIMEOUT | 30.0 | Page-load timeout (s) |
| CRAWLEMOON_API_KEY | unset | If set, every tool call must include matching _api_key |
| CRAWLEMOON_ALLOW_DANGEROUS_JS | false | Required for execute_js / execute_cdp / deobfuscate_js |
| CRAWLEMOON_JS_MAX_LENGTH | 50000 | Length cap for JS payloads |
| CRAWLEMOON_JS_EXEC_TIMEOUT | 10.0 | Per-script timeout (s) |
| CRAWLEMOON_PROXIES | unset | Comma/newline separated proxy entries |
| CRAWLEMOON_PROXIES_FILE | unset | Local file with one proxy per line |
| CRAWLEMOON_PROXY_SCHEME | http | Scheme for entries without one: http, https, socks4, socks5 |
| CRAWLEMOON_PROXY_ROTATION | round_robin | round_robin, random, sticky, least_used |
| CRAWLEMOON_PROXY_HEALTH_CHECK_INTERVAL | 300 | Proxy health-check interval (s) |
| CRAWLEMOON_PROXY_FAIL_CLOSED | false | Raise on startup proxy config errors instead of continuing direct |
Proxy, V2ray & connection control
Crawlemoon uses one rotating proxy pool for browser contexts, stealth_request, and Xray/V2ray local exits. Credentials are stored separately from normalized URLs and are masked in stats/log-style responses.
Supported proxy entry formats:
http://user:pass@31.59.20.176:6754
socks5://user:pass@127.0.0.1:1080
31.59.20.176:6754:user:pass
31.59.20.176:6754
Start the MCP server with a proxy file:
{
"mcpServers": {
"crawlemoon": {
"command": "uvx",
"args": ["crawlemoon"],
"env": {
"CRAWLEMOON_PROXIES_FILE": "/secure/path/proxies.txt",
"CRAWLEMOON_PROXY_SCHEME": "http",
"CRAWLEMOON_PROXY_ROTATION": "sticky"
}
}
}
}
Proxy files can use Webshare-style lines and comments:
# host:port:username:password
31.59.20.176:6754:uusmdewb:en3w097syrxh
31.56.127.193:7684:uusmdewb:en3w097syrxh
Configure or replace proxies at runtime with the configure_proxies MCP tool:
{
"proxies_text": "31.59.20.176:6754:user:pass\n31.56.127.193:7684:user:pass",
"default_scheme": "http",
"rotation_strategy": "sticky",
"health_check_interval": 300,
"replace_existing": true
}
Use add_proxy, remove_proxy, test_proxy, and get_proxy_stats for incremental control. get_proxy_stats returns masked proxy URLs, for example http://user:***@31.59.20.176:6754.
For V2ray/Xray, call configure_xray_subscription with either subscription_url or raw_links. Crawlemoon starts local SOCKS5 exits such as socks5://127.0.0.1:10801, registers them in the same proxy pool, and can rotate or benchmark nodes with rotate_xray_node and test_xray_nodes.
Keep proxy files out of git. They contain live credentials.
Security
execute_js, execute_cdp, and deobfuscate_js are disabled by default — they execute or operate on arbitrary code in a real browser. Enable on trusted networks with CRAWLEMOON_ALLOW_DANGEROUS_JS=true. Even then, payloads are length-capped, time-bounded, and a denylist rejects eval, new Function, dynamic import(), document.write, importScripts, and WebAssembly.{compile,instantiate}. Set CRAWLEMOON_API_KEY so MCP clients must present a matching _api_key.
These are mitigations, not a sandbox: do not expose this server to untrusted clients.
Develop
git clone https://github.com/razavioo/crawlemoon.git
cd crawlemoon
make dev-install # editable install + dev/captcha/ocr extras + pre-commit
make test # pytest
make lint # ruff + mypy
Releases
This project uses Trusted Publishing (OIDC) via GitHub Actions to automate publishing releases directly to PyPI.
To release a new version:
- Bump the version number in
pyproject.toml. - Commit the change and create a git tag matching the version (e.g.
v1.1.8):git add pyproject.toml git commit -m "chore: bump version to 1.1.8" git tag v1.1.8 - Push your branch and the tag to GitHub:
git push origin main --tags
GitHub Actions will automatically run tests, build the package, and publish it securely to PyPI under the crawlemoon package space.
PRs welcome. Particularly interested in: distributed mode (Redis queue), result sinks (Postgres / S3), Prometheus metrics. See MIT License.
Related Skills
Agent-Reach
86.1kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.1kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
ruflo
73.5k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
CowAgent
47.2kOpen-source super AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
