claude-delegator-deepseek-mcp
MCP server: delegate heavy-token tasks from Claude Code to DeepSeek, Kimi, GLM, Qwen, Grok, or any OpenAI-compatible model. Mix providers per task, cost receipt on every call. Zero dependencies.
Install / Use
claude mcp add fjgbue -- npx -y github:fjgbue/claude-delegator-deepseek-mcpIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AI & Machine LearningSupported Platforms
Tags
Our assessment of claude-delegator-deepseek-mcp
claude-delegator-deepseek-mcp scores 87/100 on our quality scale, 506th of 968 AI & Machine Learning skills we index.
Its MCP Server is 17 KB long, well organised into 19 sections with 10 code examples: a thorough specification that gives an agent plenty to work with.
It has 45 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated 16 days ago, so claude-delegator-deepseek-mcp is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 97/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful.
AI review by kimi-k2.7-code on 2026-10-10. Automated pattern scan on 2026-10-10. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
claude-delegator-deepseek-mcp compared with similar skills
All 4 of these similar skills score higher than claude-delegator-deepseek-mcp; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| claude-delegator-deepseek-mcp (this skill)by fjgbue | 87 | 45 | 16d ago | MCP Server |
| claude-memby thedotmack | 100 | 99.2k | 1d ago | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 95.5k | 2d ago | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.8k | today | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.9k | today | CLAUDE.md |
Frequently asked questions
- How do I install claude-delegator-deepseek-mcp?
- Run
claude mcp add fjgbue -- npx -y github:fjgbue/claude-delegator-deepseek-mcp. The install tabs above show the steps for each supported agent. - Which AI agents does claude-delegator-deepseek-mcp work with?
- It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
- Is claude-delegator-deepseek-mcp safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful. It is MIT-licensed and scores 97/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is claude-delegator-deepseek-mcp still maintained?
- The repository was last updated 16 days ago, so claude-delegator-deepseek-mcp is actively maintained.
Skill content
View source on GitHubClaude Code DeepSeek Delegator
🚀 3.0 is out — now model-agnostic
One tool, any model: DeepSeek, Kimi, GLM, Qwen, Grok, Groq, OpenRouter, even your local ollama — and you can mix them per task. Plus a rebuilt interactive installer, a model picker inside Claude Code, and a receipt for every cent. v2 installs keep working untouched. See what's new ↓
Cut your Claude Code bill by 90–97% on heavy work — big file audits, long generations, deep reasoning — by delegating it to a cheaper model without leaving your session.
Claude orchestrates; the delegate does the grunt work (big file audits, long generations, deep reasoning) at a fraction of the price. One tool call, no subagent spawn, no daemon, zero dependencies.
One command installs everything. init is an interactive wizard that wires up the delegate tool, the automatic gate (the "Delegate to DeepSeek? (y/n)" nudge before heavy reads and skill loads), your provider + API key (live-validated), and how models get picked. Run it once, restart Claude Code, and delegation just happens.
<img src="https://raw.githubusercontent.com/12122J/claude-delegator-deepseek-mcp/main/assets/hero.svg" alt="A Claude Code session in a light macOS terminal: the gate asks 'Delegate to DeepSeek? (y/n)', the user answers y, files are handed off via files[], a receipt line shows saved $0.194 (94% vs Opus) · spent $0.012, and Claude synthesizes the three findings" width="860">⭐ If this saves you money, please star the repo. It's the single biggest thing that helps other Claude Code users find it.
Every call ends with a receipt — shown to you automatically, straight from the API's own token counts. You always know what you spent and what you saved.
What's new in 3.0
- Any model, not just DeepSeek. 8 providers vendored (DeepSeek, Moonshot/Kimi, Z.AI + Zhipu/GLM, Alibaba/Qwen, Groq, xAI/Grok, OpenRouter) plus any OpenAI-compatible endpoint you add — including local ollama/vllm for $0 delegation.
- Mix and match providers. Select several providers in one init (space in the picker), give each its key, then route digestion to GLM-flash, code generation to DeepSeek-pro, keep a Kimi in your shortlist. Keys live in a keyring that survives re-runs and provider switches.
- A setup wizard that's actually nice. Arrow-key rail UI, live API-key validation against the real endpoint, full disclosure before anything is written, one question for model strategy.
- Pick the model your way. Smart split (cheap model digests, big model creates), a shortlist picker rendered by Claude Code's own UI, always-best, always-cheapest, or fully custom per kind of work.
- Honest receipts. Cost per call from the provider's own token counts, cached tokens billed at cached rates, savings baseline of your choice (Opus, Sonnet, or off).
- Nothing breaks. The
deepseektool name, env var, and hooks from v2 all keep working. No config = exact v2 behavior.
Install (one command)
npx claude-code-deepseek-delegator init
The wizard walks you through four choices — arrow keys, ~1 minute:
- Providers — DeepSeek, Moonshot (Kimi), Z.AI / Zhipu (GLM), Alibaba (Qwen), Groq, xAI (Grok), OpenRouter, or any custom OpenAI-compatible endpoint (ollama, vllm, LM Studio, a proxy). Enter picks one; space picks several — enroll DeepSeek and Z.AI in the same run and mix them in step 3 (the first becomes the primary that names the gate).
- API keys — one per chosen provider: detected from your environment, or paste it (hidden). The wizard live-fires a 1-token request so a bad key fails right there with the provider's real error, not on tomorrow's first delegation.
- Which model runs your delegations — a smart split (the cheap model digests, the big one creates), a shortlist you pick from each time, always best / always cheapest, or custom (see below).
- Savings baseline — measure savings against Opus 4.8, Sonnet 5, or don't show savings.
It then shows exactly what it will change, asks, applies, and prints a recap:
<img src="https://raw.githubusercontent.com/12122J/claude-delegator-deepseek-mcp/main/assets/init.svg" alt="The init wizard in a light macOS terminal: collapsed steps for Provider (DeepSeek), API key detected, live key verification, the active 'Which model runs your delegations?' picker on Smart split, the savings baseline, a panel disclosing the 4 changes, four green applied rows, the 'delegation is wired' recap panel, and the outro with a GitHub star ask" width="860">Sanity-check anytime:
npx claude-code-deepseek-delegator doctor
doctor doesn't just check that files exist — it actually fires the gate hooks, resolves every configured model against the registry, and confirms the delegation prompt is live.
What init actually changes (full disclosure)
~/.claude/delegator.json— your provider, model routing, and savings baseline. Edit it anytime; changes apply on the next call, no restart.- A clearly-labeled block in
~/.claude/CLAUDE.md— the delegation rules, fenced with<!-- >>> ... >>> -->markers. Never touches anything else in your file. - Three hooks in
~/.claude/settings.json— twoPreToolUsenudges (before large reads and skill loads) and onePostToolUsecost receipt. They only add context or display info — they never block, delete, or modify your tool calls. - An MCP server named
delegate(npx -y claude-code-deepseek-delegator) — so Claude Code shows "Calling delegate…" on every call. v2'sdeepseekserver key keeps working on untouched installs and is migrated automatically when you re-run init.
Before any of that, it writes a timestamped backup of every file it changes. To undo everything — including the config files:
npx claude-code-deepseek-delegator uninstall
Which model does the work?
Not every delegation deserves your best model. "Summarize this 2,000-line file" is digestion — a cheap model does it fine. "Rewrite this module" is creation — you want the good one. Paying pro prices for flash work is where delegation savings quietly leak.
So the wizard asks one question — "Which model runs your delegations?" — with answers that map to how people actually think:
- Smart split (recommended) — the cheap model digests big files, the big model writes code and reasons. You never think about it again; the receipt shows which one ran.
- Ask me each time — you keep a shortlist, and when you approve a delegation Claude shows it through Claude Code's native picker UI (the same one
/modeluses) with prices. You tap the model, it delegates there. - Always the best / Always the cheapest — one model for everything, zero decisions.
- Custom — pick a model per kind of work (
read/write/reason), mixing providers freely: the menus list models from every provider whose key is detected, so "reads on GLM-flash, writes on DeepSeek-pro, reasoning on Kimi" is three arrow-key picks.
What each choice feels like in a session
Smart split. You never see a model decision — the same y/n gives cheap digestion and quality creation:
❯ summarize what src/auth.py does (2,100 lines)
> Delegate to DeepSeek? (y/n) y
⎿ delegate deepseek-v4-flash via deepseek · saved $0.081 (94% vs Opus) · spent $0.005
↑ digestion → the cheap model ran
❯ now rewrite it with proper token rotation
> Delegate to DeepSeek? (y/n) y
⎿ delegate deepseek-v4-pro via deepseek · saved $0.152 (91% vs Opus) · spent $0.019
↑ creation → the big model ran
Ask me each time. After your y, Claude opens Claude Code's picker with your shortlist and prices; your tap decides:
> Delegate to DeepSeek? (y/n) y
Which model? ← Claude Code's native picker
❯ deepseek-v4-flash $0.14 / $0.28 per 1M
deepseek-v4-pro $0.435 / $0.87 per 1M
moonshot:kimi-k2.5 $0.20 / $1.20 per 1M
⎿ delegate kimi-k2.5 via moonshot · saved $0.117 (88% vs Opus) · spent $0.016
Always the best / cheapest. Every delegation is the same model — exactly how v2 behaved, one price.
Under the hood every choice just writes ~/.claude/delegator.json, which you can edit anytime — changes apply on the next call, no restart:
{
"provider": "deepseek",
"mode": "auto", // or "ask" for the shortlist picker
"shortlist": ["deepseek-v4-flash", "deepseek-v4-pro", "moonshot:kimi-k2.5"],
"routing": {
"read": "deepseek-v4-flash", // digest/summarize big inputs
"write": "deepseek-v4-pro", // generate code and docs
"reason": "deepseek-v4-pro" // math, logic, architecture
},
"baseline": "opus-4.8"
}
How the smart split works mechanically: Claude labels each delegation task: "read" | "write" | "reason", the server looks the label up in routing, and that model id goes on the wire (this chain is covered by an end-to-end test against a mock provider). An explicit model argument — like the one your picker choice produces in ask mode — always beats the routing table. If Claude omits the label entirely, the provider's default large model runs: the failure mode is a price tier, never a crash.
Providers
| Provider | id | Env var | Example models |
|----------|----|---------|----------------|
| DeepSeek | deepseek | DEEPSEEK_API_KEY | deepseek-v4-pro, deepseek-v4-flash |
| Moonshot | moonshot | MOONSHOT_API_KEY | kimi-k2.5, kimi-k2.7-code |
| Z.AI | zai | ZAI_API_KEY | glm-5, glm-4.7-flash |
| Zhipu | zhipu | ZHIPU_API_KEY | glm-5, glm-4.7 |
| Alibaba | alibaba-singapore | ALIBABA_SINGAPORE_API_KEY | qwen3.7-max, qwen3.6-flash |
| Groq | groq | GROQ_API_KEY | llama-4-maverick, kimi-k2 |
| xAI | xai | XAI_API_KEY | grok-4.5, grok-code-fast-2 |
| OpenRouter | openrouter | OPENROUTER_API_KEY | ~25 curated ids across every lab |
Provider definitions (endpoints, models, prices) follow the catwalk schema, vendored from charmbracelet's registry (MIT). Add your own in ~/.claude/delegator-providers.json — any OpenAI-compatible endpoint works, including local ones:
{
"name": "Local Ollama", "id": "ollama", "type": "openai-compat",
"api_key": "none", "api_endpoint": "http://localhost:11434/v1",
"default_large_model_id": "qwen3-coder",
"models": [{ "id": "qwen3-coder", "cost_per_1m_in": 0, "cost_per_1m_out": 0,
"context_window": 128000, "default_max_tokens": 8192 }]
}
Why this instead of spawning a subagent?
The usual pattern for heavy work — spawning a Claude subagent — starts a brand new context window: you re-pay the full context, lose your current state, and still bill at Claude rates.
This MCP server stays in your current session. Claude calls delegate(...) like any tool — no new context, no re-init, no spawn overhead — and the delegate does the heavy compute at a fraction of the rate.
The files[] trick (this is the real win)
When Claude reads files and pastes them into a prompt, those bytes land in Claude's context first — you pay Claude's rate just to pass
Truncated for display — read the full file on GitHub.
Related Skills
claude-mem
99.2kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
95.5kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Understand-Anything
85.8kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.9kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
