mcp-rag-server
Lightweight RAG server for the Model Context Protocol: ingest source code, docs, build a vector index, and expose search/citations to LLMs via MCP tools.
Install / Use
claude mcp add Daniel-Barta -- npx -y github:Daniel-Barta/mcp-rag-serverIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AI & Machine LearningSupported Platforms
Tags
Our assessment of mcp-rag-server
mcp-rag-server scores 80/100 on our quality scale, 647th of 901 AI & Machine Learning skills we index.
Its MCP Server is 25 KB long, well organised into 29 sections with 26 code examples: a thorough specification that gives an agent plenty to work with.
It has 10 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated about 5 months ago. That is recent enough to be usable, but agent tooling moves fast, so check the instructions against your agent's current version.
- Our last check on 2026-09-25 found the source still online.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 95/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful.
AI review by kimi-k2.7-code on 2026-09-24. Automated pattern scan on 2026-09-24. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
mcp-rag-server compared with similar skills
All 4 of these similar skills score higher than mcp-rag-server; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| mcp-rag-server (this skill)by Daniel-Barta | 80 | 10 | 5mo ago | MCP Server |
| claude-memby thedotmack | 100 | 95.0k | today | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 86.6k | 15d ago | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 84.9k | 2d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.2k | today | CLAUDE.md |
Frequently asked questions
- How do I install mcp-rag-server?
- Run
claude mcp add Daniel-Barta -- npx -y github:Daniel-Barta/mcp-rag-server. The install tabs above show the steps for each supported agent. - Which AI agents does mcp-rag-server work with?
- It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
- Is mcp-rag-server safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful. It is MIT-licensed and scores 95/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is mcp-rag-server still maintained?
- The repository was last updated about 5 months ago. That is recent enough to be usable, but agent tooling moves fast, so check the instructions against your agent's current version.
Skill content
View source on GitHubmcp-rag-server (RAG MCP server for any repository)
mcp-rag-server is a lightweight Retrieval‑Augmented Generation helper you can plug into any client that speaks the [Model Context Protocol (MCP)]. GitHub Copilot Agent mode in Visual Studio / VS Code is just one option – you can also use the official MCP Inspector, future MCP‑aware IDEs, or custom tooling.
It indexes a target repository directory, chunks the content (default chunk size 2400 characters, about 800 tokens, with 400 characters of overlap, about 120 tokens; both configurable via CHUNK_SIZE / CHUNK_OVERLAP), builds embeddings using either local inference via @huggingface/transformers or an OpenAI‑compatible embeddings API, and exposes MCP tools:
rag_query– semantic search returning scored snippets (path, score, snippet)read_file– secure file read (optional line range) constrained toREPO_ROOT. For PDF files, text is automatically retrieved from the unified cache file if availablelist_files– list directory contents (files & subdirectories) with optional recursion, depth and extension filtering
Two transports are supported (select with MCP_TRANSPORT=stdio|http):
stdio– simplest integration for IDEs that spawn a process (backward compatible default)http(Streamable HTTP) – recommended for large repos / first run so you can watch logs & poll readiness before attaching a client. Enable viaMCP_TRANSPORT=http. Includes DNS rebinding protection by default.
Features
- Embeddings via local inference (
@huggingface/transformers) or an OpenAI‑compatible API - Multi‑language source + docs support (configurable via
ALLOWED_EXT) - PDF support: Automatically extracts text from PDF files during indexing and caches it in a unified
pdf-text-cache.jsonfile (located alongside the index store) for fast retrieval. PDF text is treated like any other text file for semantic search - Excluded folder patterns support (configurable via
EXCLUDED_FOLDERS) - Fast glob file discovery and overlapping chunking for better recall
- Simple cosine similarity ranking (optionally swap to ANN later)
- Pluggable model selection via
MODEL_NAME(see guidance below) - Optional persistent index (multi-file storage) + warm start & incremental reindexing via
INDEX_STORE_PATH - Incremental change detection (additions / deletions / file size changes) to avoid full rebuilds
- Stdio or Streamable HTTP transport (with optional host allow‑list / DNS rebinding protection)
- Safe path handling (rejects attempts to escape
REPO_ROOT) - Minimal dependencies; quick startup after first local model load or remote API configuration validation
- Ready for extension: add new MCP tools or ANN / hybrid retrieval backends
Planned / Nice‑to‑have: hybrid BM25 + embedding search, ANN acceleration (HNSW / IVF), per‑language tokenizer heuristics, batched / parallel embedding, semantic boundary aware chunking.
Requirements
- Node.js 20+
- Path to your repository (
REPO_ROOT)
Optional MCP clients (any one is enough):
- The official MCP Inspector
- Visual Studio 2022 17.14+ with GitHub Copilot (Agent mode enabled)
- VS Code with GitHub Copilot Agent mode
- Any other MCP-aware tooling
Install
npm install
npm run build
Run (local test)
Build then start (stdio transport by default). Use either npm start or invoke the built file directly.
Windows PowerShell
npm run build
$env:REPO_ROOT="C:\path\to\your-repo"; node dist/index.js
Or:
$env:REPO_ROOT="C:\path\to\your-repo"; npm start
macOS / Linux (bash/zsh)
npm run build
export REPO_ROOT="/path/to/your-repo"; node dist/index.js
Or:
export REPO_ROOT="/path/to/your-repo"; npm start
Optionally set a model cache to speed up subsequent runs (first start downloads the model once):
export TRANSFORMERS_CACHE="/path/to/cache" # macOS/Linux
$env:TRANSFORMERS_CACHE="C:\path\to\cache" # Windows PowerShell
OpenAI-compatible API embeddings
Set EMBEDDING_PROVIDER=openai to call a remote /embeddings endpoint instead of loading a local transformer model. The request format follows the OpenAI embeddings API and works with providers that expose a compatible protocol such as OpenAI, Mistral, and Jina AI.
Windows PowerShell:
$env:REPO_ROOT="C:\path\to\your-repo"
$env:EMBEDDING_PROVIDER="openai"
$env:EMBEDDING_API_BASE_URL="https://api.openai.com/v1"
$env:EMBEDDING_API_KEY="<your-api-key>"
$env:MODEL_NAME="text-embedding-3-small"
npm start
macOS / Linux:
export REPO_ROOT="/path/to/your-repo"
export EMBEDDING_PROVIDER="openai"
export EMBEDDING_API_BASE_URL="https://api.openai.com/v1"
export EMBEDDING_API_KEY="<your-api-key>"
export MODEL_NAME="text-embedding-3-small"
npm start
Notes:
EMBEDDING_API_BASE_URLshould point to the provider's API base (for examplehttps://api.openai.com/v1), not the/embeddingspath itself.MODEL_NAMEis passed verbatim to the remote embeddings API whenEMBEDDING_PROVIDER=openai.EMBEDDING_API_BATCH_SIZEcontrols how many chunks are sent per remote embeddings request during indexing. Default:200.TRANSFORMERS_CACHEis only relevant for local inference.
Streamable HTTP mode (recommended for large initial indexes)
Run the MCP server as an HTTP endpoint and only open your IDE after Embeddings ready. shows (avoids client timeouts on cold start):
npm run build
$env:REPO_ROOT="C:\path\to\your-repo"; $env:MCP_TRANSPORT="http"; npm start
export REPO_ROOT="/path/to/your-repo"; MCP_TRANSPORT=http npm start
Default HTTP bind: http://127.0.0.1:3000/mcp. Override with HOST and MCP_PORT envs. A readiness endpoint is available at http://127.0.0.1:3000/health returning JSON like:
{
"version": "0.x.y",
"repoRoot": "C:/abs/path",
"modelName": "<embedding model>",
"transport": "stdio" | "http",
"ready": true | false,
"startedAt": "2025-01-01T00:00:00.000Z",
"indexing": {
"filesDiscovered": 123,
"chunksTotal": 456,
"chunksEmbedded": 456
}
}
ready flips to true only once all discovered chunks have embeddings (post cold build or incremental update completion).
Instructions endpoint
The server also exposes GET /instructions, which serves the Markdown file docs/copilot-instructions.md with all occurrences of <FOLDER_INFO_NAME> replaced by the FOLDER_INFO_NAME value from your environment (default REPO_ROOT).
Notes:
- Start the server from the repository root so
docs/copilot-instructions.mdresolves via the current working directory. - Response content type is
text/markdown; charset=utf-8.
Linting & Formatting
- Type-check (no emit):
npm run typecheck - Run ESLint (check):
npm run lint - Auto-fix ESLint issues:
npm run lint:fix - Format with Prettier:
npm run format - Check formatting:
npm run format:check
Test with MCP Inspector (without VS)
Use the MCP Inspector to exercise the server locally and try the tools without Visual Studio.
Windows PowerShell:
npm run build
$env:REPO_ROOT="C:\path\to\your-repo"; npx @modelcontextprotocol/inspector node .\\dist\\index.js
Streamable HTTP via Inspector (Windows):
npm run build
$env:REPO_ROOT="C:\path\to\your-repo"; $env:MCP_TRANSPORT="http"; npx @modelcontextprotocol/inspector http://localhost:3000/mcp --transport http
macOS/Linux (bash/zsh):
export REPO_ROOT="/path/to/your-repo"
npx @modelcontextprotocol/inspector node dist/index.js
Streamable HTTP (macOS/Linux):
export REPO_ROOT="/path/to/your-repo"; MCP_TRANSPORT=http npx @modelcontextprotocol/inspector http://localhost:3000/mcp --transport http
Notes:
- First run in local mode downloads the embedding model and builds embeddings; the Inspector will connect only after startup completes. Watch the terminal for progress logs printed to stderr.
- You can also put settings in a
.envfile at the project root (e.g.,REPO_ROOT,TRANSFORMERS_CACHE,EMBEDDING_PROVIDER,EMBEDDING_API_BASE_URL).
In the Inspector UI:
- Click "List tools" to verify these tools are available:
rag_query,read_file,list_files. - Select a tool and click "Call tool". Provide JSON input as shown below.
Examples
- Semantic search over the repo
Tool: rag_query
Input JSON:
{
"query": "protobuf message X schema",
"top_k": 5
}
The response includes an array of matches with path, score, and snippet.
- List files in a directory (non-recursive by default)
Tool: list_files
Input JSON:
{
"dir": "src",
"recursive": false
}
Recursive with filters and limits:
Tool: list_files
Input JSON:
{
"dir": "src",
"recursive": true,
"maxDepth": 3,
"includeExtensions": ["ts", "md"],
"limit": 200
}
Response shape:
{
"entries": [
{ "path": "src/", "type": "dir" },
{ "path": "src/index.ts", "type": "file", "size": 1234 },
{ "path": "src/lib/", "type": "dir" }
]
}
- Read a file (optionally with a line range)
Tool: read_file
Input JSON:
{
"path": "src/path/to/file.txt", // relative to REPO_ROOT
"startLine": 1,
"endLine": 120
}
Troubleshooting
- Slow startup: set
TRANSFORMERS_CACHEto a fast local folder and (optionally) setALLOWED_EXT(e.g.,ts,tsx,jsfor TypeScript/JS only, or any list you need). - Path errors:
pathmust be relative toREPO_ROOT. Absolute paths are rejected for safety. - Nothing appears in Inspector for minutes: the server is still initializing (model download + embedding). This is expected on first run.
- Slow warm restarts: provide
INDEX_STORE_PATHso embeddings persist and only changed files re‑embed.
Environment configuration (.env)
You can configure environment variables via a local .env file.
Steps:
- Copy
.env.exampleto.env. - Edit values as needed.
Supported variables:
REPO_ROOT(required): path to the repository to index.FOLDER_INFO_NAME(optional): display label used inside MCP tool descriptions for the repository root (defaultREPO_ROOT). This is purely cosmetic for client UX; it does NOT affect which directory is indexed (that is controlled only byREPO_ROOT). Set it if you prefer a friendlier name (e.g.,frontend-appormonorepo-root) to appear in tool metadata and path guidance returned to the client.EMBEDDING_PROVIDER(optional):local(default) oropenai.openaimeans “use an OpenAI-compatible/embeddingsAPI”, not specifically OpenAI as the vendor.TRANSFORMERS_CACHE(optional): cache folder for local model files.EMBEDDING_API_BASE_URL(required whenEMBEDDING_PROVIDER=openai): base URL for the OpenAI-compatible API, such ashttps://api.openai.com/v1,https://api.mistral.ai/v1, or your provider-specific equivalent.EMBEDDING_API_KEY(required whenEMBEDDING_PROVIDER=openai): bearer token used for the embeddings API.EMBEDDING_API_BATCH_SIZE(optional whenEMBEDDING_PROVIDER=openai): number of chunks sent per remote embeddings request during indexing. Default:200.ALLOWED_EXT(optional): comma-separated list of file extensions to index. Default includes common text/code formats pluspdf. PDF files are automatically processed: text is extracted once during indexing and cached in a unifiedpdf-text-cache.jsonfile for fast retrieval.EXCLUDED_FOLDERS(optional): comma-separated list of folder patterns to exclude from indexing. Supports both exact folder names (e.g.,node_modules,dist,build,.git) and basic glob patterns (e.g.,**/test/**,**/tests/**). Files in these folders will be skipped during indexing. Defaults include common build/dependency folders:node_modules,dist,build,.git,target,bin,obj,.cache,coverage,.nyc_output.MCP_TRANSPORT(optional):httporstdio.VERBOSE(optional): true/1/yes/on for more
Truncated for display — read the full file on GitHub.
Related Skills
claude-mem
95.0kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
86.6kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Understand-Anything
84.9kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.2kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
