SkillAgentSearch skills...

omnifeed

LLM-friendly web crawler & scraper with a dedicated Reddit engine, built on Crawl4AI — Open WebUI compatible

Install / Use

claude mcp add kinorai -- npx -y github:kinorai/omnifeed

If the server publishes to npm under a different name, use that package instead — check the repo README.

About this skill
🔌

MCP Server

Model Context Protocol server

Quality Score

81/100

Supported Platforms

Claude Code
Claude Desktop
OpenAI Codex

Tags

Our assessment of omnifeed

omnifeed scores 81/100 on our quality scale, 731st of 956 AI & Machine Learning skills we index.

Its MCP Server is 14 KB long, well organised into 12 sections with 7 code examples: a thorough specification that gives an agent plenty to work with.

It has 10 GitHub stars, so there is little community track record yet; judge it on its content.

Substance
30/30
Structure
20/20
Description
12/15
Adoption
4/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated today, so omnifeed is actively maintained.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 97/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

omnifeed compared with similar skills

All 4 of these similar skills score higher than omnifeed; compare them before choosing.

SkillScoreStarsUpdatedFormat
omnifeed (this skill)by kinorai8110todayMCP Server
claude-memby thedotmack10097.9k1d agoCLAUDE.md
Agent-Reachby Panniantong10093.6ktodayCLAUDE.md
Understand-Anythingby Egonex-AI10085.6k2d agoCLAUDE.md
headroomby headroomlabs-ai10074.6ktodayCLAUDE.md

Frequently asked questions

How do I install omnifeed?
Run claude mcp add kinorai -- npx -y github:kinorai/omnifeed. The install tabs above show the steps for each supported agent.
Which AI agents does omnifeed work with?
It is written for Claude Code, Claude Desktop and OpenAI Codex, as a MCP Server file. Other agents that read the same format can often use it too.
Is omnifeed safe to use?
It is MIT-licensed and scores 97/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is omnifeed still maintained?
The repository was last updated today, so omnifeed is actively maintained.
<!-- markdownlint-disable MD033 MD041 --> <p align="center"> <img src="https://capsule-render.vercel.app/api?type=waving&color=0:FF4500,100:7C3AED&height=220&section=header&text=omnifeed&fontSize=82&fontColor=ffffff&animation=fadeIn&fontAlignY=36" alt="omnifeed" width="100%"/> </p> <p align="center"> <img src="https://readme-typing-svg.demolab.com?font=Fira+Code&weight=700&size=28&color=FF4500&center=true&vCenter=true&multiline=true&repeat=false&duration=1500&pause=500&width=860&height=110&lines=Self-hosted+web+search+%2B+fetch+MCP;with+a+dedicated+Reddit+engine+%E2%80%94+and+more" alt="Self-hosted web search + fetch MCP, with a dedicated Reddit engine — and more"/> </p> <p align="center"> <a href="https://github.com/kinorai/omnifeed/actions/workflows/ci.yml"><img src="https://img.shields.io/github/actions/workflow/status/kinorai/omnifeed/ci.yml?branch=main&label=CI&style=flat-square" alt="CI"/></a> <a href="https://github.com/kinorai/omnifeed/releases"><img src="https://img.shields.io/github/v/release/kinorai/omnifeed?style=flat-square&color=FF4500" alt="Release"/></a> <a href="https://hub.docker.com/r/kinorai/omnifeed"><img src="https://img.shields.io/docker/pulls/kinorai/omnifeed?style=flat-square&logo=docker&logoColor=white&color=2496ED" alt="Docker pulls"/></a> </p> <p align="center"> omnifeed gives an AI agent the full research loop, <b>search → URLs → content</b>, on self-hosted <a href="https://github.com/searxng/searxng">SearXNG</a> and <a href="https://github.com/unclecode/crawl4ai">crawl4ai</a>. Its <b>Reddit engine</b> returns full comment trees as <a href="https://github.com/toon-format/toon">TOON</a>, lossless and about 40% fewer tokens than JSON, with <b>no Reddit API key</b>. Hacker News, GitHub, Bluesky, Discourse and X/Twitter get their own engines too. </p>
  • web_search queries SearXNG (Google, Bing, DDG, Reddit included) and returns ranked URLs with titles and snippets. Pass site to scope results to one hostname. Naming the site in the query text fails, because engines read it as a topic word.
  • fetch_url returns any URL as clean markdown through crawl4ai. Dedicated engines return TOON instead: Reddit threads and /r/{sub} listings through a real browser (listings honor the URL's ?t= and ?limit=), plus Hacker News, GitHub issues and pull requests, Bluesky posts and profiles, and Discourse topics from their public APIs. The GitHub engine also returns compact markdown for repository roots (metadata, latest release, README), files, directories, releases, commits, gists and, with a token, discussions. X/Twitter post links (x.com, twitter.com, the fxtwitter/vxtwitter mirrors and t.co) return markdown with the full text, quote, media alt text, community note, poll, the author's thread and top replies, read from FxTwitter with syndication, vxTwitter and crawl4ai as fallbacks.
<img src="https://user-images.githubusercontent.com/74038190/212284100-561aa473-3905-4a80-b561-0d28506553ee.gif" width="100%">

<img src="https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Activities/Sparkles.png" width="26" height="26" /> Why omnifeed

| | omnifeed | Cloud web MCPs and other Reddit MCPs | |---|---|---| | Works on Reddit | ✅ your residential IP + real browser | ❌ datacenter IPs get 403 | | Search and crawl in one self-hosted service | ✅ SearXNG + crawl4ai | ❌ search-only or crawl-only | | Full comment tree (/api/morechildren expansion) | ✅ up to 40 rounds, about 4k comments | ❌ | | Token-efficient output | ✅ TOON, about 40% smaller than JSON | ❌ verbose JSON or truncated bodies | | Generic crawl for non-Reddit URLs | ✅ via crawl4ai | ❌ | | Front-ends | MCP, Open WebUI, REST | MCP only, mostly |

<img src="https://user-images.githubusercontent.com/74038190/212284100-561aa473-3905-4a80-b561-0d28506553ee.gif" width="100%">

<img src="https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Travel%20and%20places/Rocket.png" width="26" height="26" /> Quick start

# Fetch the compose file and SearXNG settings, then start:
curl -fsSL https://raw.githubusercontent.com/kinorai/omnifeed/main/docker-compose.yml -o docker-compose.yml
curl -fsSL --create-dirs https://raw.githubusercontent.com/kinorai/omnifeed/main/searxng/settings.yml -o searxng/settings.yml
docker compose up

This starts omnifeed, SearXNG and crawl4ai, <b>tokenless out of the box</b>, because the compose file sets OMNIFEED_DEV_NO_AUTH=true. searxng/settings.yml enables the json format that web_search needs. For Open WebUI, set WEB_LOADER_ENGINE=external and point it at http://localhost:8080. To require a token, see Authentication below.

On Apple Silicon you can skip Docker and use Apple's native container runtime. See docs/apple-container.md.

<img src="https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Objects/Electric%20Plug.png" width="22" height="22" /> As an MCP server

Works with any MCP client, including Claude Code, Cursor, Codex, Gemini CLI, OpenCode, Windsurf and Pi. One endpoint serves the stateless MCP protocol and the older initialize-era revisions. Stateless requests get the spec's HTTP statuses, 400 for header or version violations and 404 for unknown methods. Initialize-era responses stay 200. omnifeed rejects cross-origin browser requests unless OMNIFEED_ALLOWED_ORIGINS lists them.

HTTP, recommended. docker compose up already serves MCP on :8081. Point your client at it:

{ "mcpServers": { "omnifeed": { "url": "http://localhost:8081/mcp" } } }

Stdio, for clients that speak nothing else. The client spawns and owns a stdio server, so it can't be a long-running compose service. Launch the compose file's mcp profile instead, which reuses its upstreams, network and image:

{
  "mcpServers": {
    "omnifeed": {
      "command": "docker",
      "args": ["compose", "-f", "/abs/path/to/docker-compose.yml", "run", "-T", "--rm", "mcp"]
    }
  }
}

run -T disables the TTY so JSON-RPC pipes cleanly. Start the stack first with docker compose up -d so the upstreams are healthy.

<details> <summary><b>Standalone stdio, without the compose stack</b></summary>

Spawn the container and tell it where crawl4ai and SearXNG are. omnifeed exits at startup without OMNIFEED_CRAWL4AI_URL.

{
  "mcpServers": {
    "omnifeed": {
      "command": "docker",
      "args": [
        "run", "--rm", "-i",
        "-e", "OMNIFEED_CRAWL4AI_URL=http://host.docker.internal:11235/crawl",
        "-e", "OMNIFEED_SEARXNG_URL=http://host.docker.internal:8080",
        "kinorai/omnifeed:latest", "--mcp-stdio"
      ]
    }
  }
}

On Linux, add "--add-host=host.docker.internal:host-gateway" to the args.

</details>

fetch_url is always available. web_search appears only when OMNIFEED_SEARXNG_URL is set. The agent calls web_search, picks URLs, then calls fetch_url.

/crawl returns [{"page_content": "...", "metadata": {...}}], the shape of a LangChain or LlamaIndex Document, so a custom document loader takes a few lines.

<img src="https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Objects/Locked%20with%20Key.png" width="22" height="22" /> Authentication

The compose stack runs tokenless for local use. To require a bearer token, generate one:

openssl rand -hex 32

Set OMNIFEED_API_KEY to it in docker-compose.yml and remove OMNIFEED_DEV_NO_AUTH. Clients send it as Authorization: Bearer <token>. With neither variable set, omnifeed refuses to start, so it can't be left open by accident. Stdio MCP needs no token. It inherits the trust of the process that spawned it.

<img src="https://user-images.githubusercontent.com/74038190/212284100-561aa473-3905-4a80-b561-0d28506553ee.gif" width="100%">

<img src="https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Objects/Gear.png" width="26" height="26" /> Configuration

omnifeed reads OMNIFEED_-prefixed environment variables. You usually set only three: OMNIFEED_API_KEY, OMNIFEED_CRAWL4AI_URL and, for search, OMNIFEED_SEARXNG_URL.

docs/configuration.md lists every variable, plus fetch truncation (max_chars and start_char), infinite-scroll fetching, Reddit size limits, the response cache (no_cache to bypass), partial Reddit threads, per-engine timeouts and Prometheus metrics.

Running more than one replica? Set OMNIFEED_REDIS_URL so the rate limiters share state and the deployment obeys one limit. Without it, N replicas send N times the configured rate, which upstream search engines notice. If Redis goes down, the limiters fall back to per-process pacing and crawls keep working.

<img src="https://user-images.githubusercontent.com/74038190/212284100-561aa473-3905-4a80-b561-0d28506553ee.gif" width="100%">

<img src="https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Travel%20and%20places/Building%20Construction.png" width="26" height="26" /> Architecture

%%{init: {"theme":"base","themeVariables":{"background":"transparent","mainBkg":"#161b22","primaryColor":"#161b22","primaryTextColor":"#e6edf3","primaryBorderColor":"#FF4500","lineColor":"#8b949e","secondaryColor":"#161b22","tertiaryColor":"#161b22"},"flowchart":{"curve":"basis","htmlLabels":false}}}%%
flowchart TB
  crawl["POST /crawl"] e1@--> owt["Open WebUI<br/>transport"]
  search["POST /search"] e2@--> sat["SearchAPI<br/>transport"]
  mcpStdio["MCP stdio"] e3@--> mcp["MCP server"]
  mcpHTTP["MCP HTTP /mcp"] e4@--> mcp

  owt e5@--> reg["Engine Registry"]
  mcp -- crawl tools --> reg
  sat e6@--> searcher["Searcher<br/>(SearXNG)"]
  mcp -- search tool --> searcher

  reg e7@--> reddit["Reddit engine<br/>(TOON)"]
  reg e12@--> hn["Hacker News engine<br/>(TOON)"]
  reg e14@--> gh["GitHub engine<br/>(TOON + markdown)"]
  reg e16@--> disc["Discourse engine<br/>(TOON)"]
  reg e8@--> generic["Generic fallback<br/>(markdown)"]
  reddit e9@--> c4["crawl4ai upstream<br/>(headless browser)"]
  generic e10@--> c4
  hn e13@--> algolia["Algolia HN API<br/>(hn.algolia.com)"]
  gh e15@--> ghapi["GitHub REST + GraphQL API<br/>(api.github.com)"]
  disc e17@--> discapi["Discourse topic JSON<br/>(allowlisted forums)"]
  searcher e11@--> sx["SearXNG upstream<br/>(Google / Bing / DDG)"]

  classDef box fill:#161b22,stroke:#30363d,stroke-width:1px,color:#e6edf3;
  classDef accent fill:#0d1117,stroke:#FF4500,stroke-width:2px,color:#ffd9b3;
  classDef animate stroke:#FF4500,stroke-width:2px,stroke-dasharray:10 6,stroke-dashoffset:900,animation:dash 14s linear infinite;
  class crawl,search,mcpStdio,mcpHTTP,owt,sat box;
  class mcp,reg,searcher,reddit,hn,gh,disc,generic,c4,sx,algolia,ghapi,discapi accent;
  class e1,e2,e3,e4,e5,e6,e7,e8,e9,e10,e11,e12,e13,e14,e15,e16,e17 animate;

<img src="https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Objects/Shield.png" width="22" height="22" /> Reddit anti-bot handling

Reddit 403-blocks non-browser HTTP clients. The Reddit engine drives a real headless browser to a www.reddit.com page and fetches Reddit's JSON from inside it, with no auth, cookies or API key. Sustained scraping can still raise your IP's risk score, so slow down if fetches return the block page. Details and tuning.

<img src="https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Activities/Puzzle%20Piece.png" width="22" height="22" /> Extending it

Engines, searchers, MCP tools and transports each plug into one small

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars10
CategoryAI
Updated2h ago
Forks2

Languages

Go

Trust signals

97/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 info