omnifeed
LLM-friendly web crawler & scraper with a dedicated Reddit engine, built on Crawl4AI — Open WebUI compatible
Install / Use
claude mcp add kinorai -- npx -y github:kinorai/omnifeedIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AI & Machine LearningSupported Platforms
Our assessment of omnifeed
omnifeed scores 81/100 on our quality scale, 731st of 956 AI & Machine Learning skills we index.
Its MCP Server is 14 KB long, well organised into 12 sections with 7 code examples: a thorough specification that gives an agent plenty to work with.
It has 10 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated today, so omnifeed is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 97/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
omnifeed compared with similar skills
All 4 of these similar skills score higher than omnifeed; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| omnifeed (this skill)by kinorai | 81 | 10 | today | MCP Server |
| claude-memby thedotmack | 100 | 97.9k | 1d ago | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 93.6k | today | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.6k | 2d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.6k | today | CLAUDE.md |
Frequently asked questions
- How do I install omnifeed?
- Run
claude mcp add kinorai -- npx -y github:kinorai/omnifeed. The install tabs above show the steps for each supported agent. - Which AI agents does omnifeed work with?
- It is written for Claude Code, Claude Desktop and OpenAI Codex, as a MCP Server file. Other agents that read the same format can often use it too.
- Is omnifeed safe to use?
- It is MIT-licensed and scores 97/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is omnifeed still maintained?
- The repository was last updated today, so omnifeed is actively maintained.
Skill content
View source on GitHubweb_searchqueries SearXNG (Google, Bing, DDG, Reddit included) and returns ranked URLs with titles and snippets. Passsiteto scope results to one hostname. Naming the site in the query text fails, because engines read it as a topic word.fetch_urlreturns any URL as clean markdown through crawl4ai. Dedicated engines return TOON instead: Reddit threads and/r/{sub}listings through a real browser (listings honor the URL's?t=and?limit=), plus Hacker News, GitHub issues and pull requests, Bluesky posts and profiles, and Discourse topics from their public APIs. The GitHub engine also returns compact markdown for repository roots (metadata, latest release, README), files, directories, releases, commits, gists and, with a token, discussions. X/Twitter post links (x.com,twitter.com, the fxtwitter/vxtwitter mirrors andt.co) return markdown with the full text, quote, media alt text, community note, poll, the author's thread and top replies, read from FxTwitter with syndication, vxTwitter and crawl4ai as fallbacks.
<img src="https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Activities/Sparkles.png" width="26" height="26" /> Why omnifeed
| | omnifeed | Cloud web MCPs and other Reddit MCPs |
|---|---|---|
| Works on Reddit | ✅ your residential IP + real browser | ❌ datacenter IPs get 403 |
| Search and crawl in one self-hosted service | ✅ SearXNG + crawl4ai | ❌ search-only or crawl-only |
| Full comment tree (/api/morechildren expansion) | ✅ up to 40 rounds, about 4k comments | ❌ |
| Token-efficient output | ✅ TOON, about 40% smaller than JSON | ❌ verbose JSON or truncated bodies |
| Generic crawl for non-Reddit URLs | ✅ via crawl4ai | ❌ |
| Front-ends | MCP, Open WebUI, REST | MCP only, mostly |
<img src="https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Travel%20and%20places/Rocket.png" width="26" height="26" /> Quick start
# Fetch the compose file and SearXNG settings, then start:
curl -fsSL https://raw.githubusercontent.com/kinorai/omnifeed/main/docker-compose.yml -o docker-compose.yml
curl -fsSL --create-dirs https://raw.githubusercontent.com/kinorai/omnifeed/main/searxng/settings.yml -o searxng/settings.yml
docker compose up
This starts omnifeed, SearXNG and crawl4ai, <b>tokenless out of the box</b>, because the compose file sets OMNIFEED_DEV_NO_AUTH=true. searxng/settings.yml enables the json format that web_search needs. For Open WebUI, set WEB_LOADER_ENGINE=external and point it at http://localhost:8080. To require a token, see Authentication below.
On Apple Silicon you can skip Docker and use Apple's native container runtime. See docs/apple-container.md.
<img src="https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Objects/Electric%20Plug.png" width="22" height="22" /> As an MCP server
Works with any MCP client, including Claude Code, Cursor, Codex, Gemini CLI, OpenCode, Windsurf and Pi. One endpoint serves the stateless MCP protocol and the older initialize-era revisions. Stateless requests get the spec's HTTP statuses, 400 for header or version violations and 404 for unknown methods. Initialize-era responses stay 200. omnifeed rejects cross-origin browser requests unless OMNIFEED_ALLOWED_ORIGINS lists them.
HTTP, recommended. docker compose up already serves MCP on :8081. Point your client at it:
{ "mcpServers": { "omnifeed": { "url": "http://localhost:8081/mcp" } } }
Stdio, for clients that speak nothing else. The client spawns and owns a stdio server, so it can't be a long-running compose service. Launch the compose file's mcp profile instead, which reuses its upstreams, network and image:
{
"mcpServers": {
"omnifeed": {
"command": "docker",
"args": ["compose", "-f", "/abs/path/to/docker-compose.yml", "run", "-T", "--rm", "mcp"]
}
}
}
run -T disables the TTY so JSON-RPC pipes cleanly. Start the stack first with docker compose up -d so the upstreams are healthy.
Spawn the container and tell it where crawl4ai and SearXNG are. omnifeed exits at startup without OMNIFEED_CRAWL4AI_URL.
{
"mcpServers": {
"omnifeed": {
"command": "docker",
"args": [
"run", "--rm", "-i",
"-e", "OMNIFEED_CRAWL4AI_URL=http://host.docker.internal:11235/crawl",
"-e", "OMNIFEED_SEARXNG_URL=http://host.docker.internal:8080",
"kinorai/omnifeed:latest", "--mcp-stdio"
]
}
}
}
On Linux, add "--add-host=host.docker.internal:host-gateway" to the args.
fetch_url is always available. web_search appears only when OMNIFEED_SEARXNG_URL is set. The agent calls web_search, picks URLs, then calls fetch_url.
/crawl returns [{"page_content": "...", "metadata": {...}}], the shape of a LangChain or LlamaIndex Document, so a custom document loader takes a few lines.
<img src="https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Objects/Locked%20with%20Key.png" width="22" height="22" /> Authentication
The compose stack runs tokenless for local use. To require a bearer token, generate one:
openssl rand -hex 32
Set OMNIFEED_API_KEY to it in docker-compose.yml and remove OMNIFEED_DEV_NO_AUTH. Clients send it as Authorization: Bearer <token>. With neither variable set, omnifeed refuses to start, so it can't be left open by accident. Stdio MCP needs no token. It inherits the trust of the process that spawned it.
<img src="https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Objects/Gear.png" width="26" height="26" /> Configuration
omnifeed reads OMNIFEED_-prefixed environment variables. You usually set only three: OMNIFEED_API_KEY, OMNIFEED_CRAWL4AI_URL and, for search, OMNIFEED_SEARXNG_URL.
docs/configuration.md lists every variable, plus fetch truncation (max_chars and start_char), infinite-scroll fetching, Reddit size limits, the response cache (no_cache to bypass), partial Reddit threads, per-engine timeouts and Prometheus metrics.
Running more than one replica? Set OMNIFEED_REDIS_URL so the rate limiters share state and the deployment obeys one limit. Without it, N replicas send N times the configured rate, which upstream search engines notice. If Redis goes down, the limiters fall back to per-process pacing and crawls keep working.
<img src="https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Travel%20and%20places/Building%20Construction.png" width="26" height="26" /> Architecture
%%{init: {"theme":"base","themeVariables":{"background":"transparent","mainBkg":"#161b22","primaryColor":"#161b22","primaryTextColor":"#e6edf3","primaryBorderColor":"#FF4500","lineColor":"#8b949e","secondaryColor":"#161b22","tertiaryColor":"#161b22"},"flowchart":{"curve":"basis","htmlLabels":false}}}%%
flowchart TB
crawl["POST /crawl"] e1@--> owt["Open WebUI<br/>transport"]
search["POST /search"] e2@--> sat["SearchAPI<br/>transport"]
mcpStdio["MCP stdio"] e3@--> mcp["MCP server"]
mcpHTTP["MCP HTTP /mcp"] e4@--> mcp
owt e5@--> reg["Engine Registry"]
mcp -- crawl tools --> reg
sat e6@--> searcher["Searcher<br/>(SearXNG)"]
mcp -- search tool --> searcher
reg e7@--> reddit["Reddit engine<br/>(TOON)"]
reg e12@--> hn["Hacker News engine<br/>(TOON)"]
reg e14@--> gh["GitHub engine<br/>(TOON + markdown)"]
reg e16@--> disc["Discourse engine<br/>(TOON)"]
reg e8@--> generic["Generic fallback<br/>(markdown)"]
reddit e9@--> c4["crawl4ai upstream<br/>(headless browser)"]
generic e10@--> c4
hn e13@--> algolia["Algolia HN API<br/>(hn.algolia.com)"]
gh e15@--> ghapi["GitHub REST + GraphQL API<br/>(api.github.com)"]
disc e17@--> discapi["Discourse topic JSON<br/>(allowlisted forums)"]
searcher e11@--> sx["SearXNG upstream<br/>(Google / Bing / DDG)"]
classDef box fill:#161b22,stroke:#30363d,stroke-width:1px,color:#e6edf3;
classDef accent fill:#0d1117,stroke:#FF4500,stroke-width:2px,color:#ffd9b3;
classDef animate stroke:#FF4500,stroke-width:2px,stroke-dasharray:10 6,stroke-dashoffset:900,animation:dash 14s linear infinite;
class crawl,search,mcpStdio,mcpHTTP,owt,sat box;
class mcp,reg,searcher,reddit,hn,gh,disc,generic,c4,sx,algolia,ghapi,discapi accent;
class e1,e2,e3,e4,e5,e6,e7,e8,e9,e10,e11,e12,e13,e14,e15,e16,e17 animate;
<img src="https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Objects/Shield.png" width="22" height="22" /> Reddit anti-bot handling
Reddit 403-blocks non-browser HTTP clients. The Reddit engine drives a real headless browser to a www.reddit.com page and fetches Reddit's JSON from inside it, with no auth, cookies or API key. Sustained scraping can still raise your IP's risk score, so slow down if fetches return the block page. Details and tuning.
<img src="https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Activities/Puzzle%20Piece.png" width="22" height="22" /> Extending it
Engines, searchers, MCP tools and transports each plug into one small
Truncated for display — read the full file on GitHub.
Related Skills
claude-mem
97.9kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
93.6kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Understand-Anything
85.6kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.6kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
