spellbook
Multi-platform AI assistant skills and workflows. Serious engineering. Also fun.
Install / Use
claude mcp add axiomantic -- npx -y github:axiomantic/spellbookIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AutomationSupported Platforms
Our assessment of spellbook
spellbook scores 72/100 on our quality scale, 1893rd of 2,176 Automation skills we index.
Its MCP Server is 64 KB long, well organised into 45 sections with 14 code examples: long enough that it reads more like full documentation than a focused instruction file, which agents can find harder to follow.
It has 10 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated 6 days ago, so spellbook is actively maintained.
- Our last check on 2026-09-24 found the source still online.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 92/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
ReviewOur scan of the whole file found 3 patterns worth reviewing before you install spellbook. An AI review of the same text found nothing harmful.
- mediumDisables the agent's permission promptsline 56
- [YOLO Mode](#yolo-mode) - mediumDisables the agent's permission promptsline 554
### YOLO Mode - mediumDisables the agent's permission promptsline 557
> **YOLO mode gives your AI assistant full control of your system.** - noteInstalls by piping a downloaded script into a shellline 83
curl -fsSL https://raw.githubusercontent.com/axiomantic/spellbook/main/bootstrap.sh | bash - noteInstalls by piping a downloaded script into a shellline 97
irm https://raw.githubusercontent.com/axiomantic/spellbook/main/bootstrap.ps1 | iex
AI review by kimi-k2.7-code on 2026-09-24. Automated pattern scan on 2026-09-24. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
spellbook compared with similar skills
All 4 of these similar skills score higher than spellbook; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| spellbook (this skill)by axiomantic | 72 | 10 | 6d ago | MCP Server |
| Agent-Reachby Panniantong | 100 | 86.1k | 13d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.1k | today | CLAUDE.md |
| rufloby ruvnet | 100 | 73.5k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.2k | today | CLAUDE.md |
Frequently asked questions
- How do I install spellbook?
- Run
claude mcp add axiomantic -- npx -y github:axiomantic/spellbook. The install tabs above show the steps for each supported agent. - Which AI agents does spellbook work with?
- It is written for Claude Code, Claude Desktop, Gemini CLI and OpenAI Codex, as a MCP Server file. Other agents that read the same format can often use it too.
- Is spellbook safe to use?
- Our scan of the whole file found 3 patterns worth reviewing before you install spellbook. An AI review of the same text found nothing harmful. It is MIT-licensed and scores 92/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is spellbook still maintained?
- The repository was last updated 6 days ago, so spellbook is actively maintained.
Skill content
View source on GitHubTable of Contents
<!-- START doctoc generated TOC please keep comment here to allow auto update --> <!-- DON'T EDIT THIS SECTION, INSTEAD RE-RUN doctoc TO UPDATE -->- Quick Install
- What Spellbook Does
- The develop Skill
- What's Included
- Platform Support
- Example Workflows
- Recommended Companion Tools
- Key Skills
- Development
- Research backing
- Documentation
- Contributing
- Acknowledgments
- Attribution
- License
Quick Install
curl -fsSL https://raw.githubusercontent.com/axiomantic/spellbook/main/bootstrap.sh | bash
The installer requires Python 3.10+ and git, then automatically installs uv and configures skills for detected platforms.
Upgrade: cd ~/.local/share/spellbook && git pull && python3 install.py
Uninstall: python3 ~/.local/share/spellbook/uninstall.py
See Installation Guide for advanced options.
Windows Quickstart
irm https://raw.githubusercontent.com/axiomantic/spellbook/main/bootstrap.ps1 | iex
Requirements: Python 3.10+, git, and PowerShell 5.1+.
- Symlinks require Developer Mode enabled in Windows Settings (falls back to junctions or copies otherwise)
- Service management uses Windows Task Scheduler
- Install location:
%LOCALAPPDATA%\spellbook
What Spellbook Does
Spellbook is a harness-augmentation layer for AI coding assistants. The harness is the runtime that hosts the agent loop and executes tools (Claude Code, Codex, OpenCode, Gemini CLI, ForgeCode). Spellbook plugs into whichever harness you are running and adds skills, slash commands, hooks, profiles, and a shared MCP server (memory, focus stints, session resume) on top.
Three things distinguish it from harness-native features and from other skill collections:
- Harness-agnostic. The same skills, commands, and memory work across every supported harness on the same project. Switch from Claude Code to OpenCode mid-task and the workflow continues.
- Shared centralized MCP server. Memories, focus stints, and session-resume state live in one place, so context stored from a Claude Code session surfaces in an OpenCode session on the same repo. No individual harness ships this.
- Skills + hooks layer no harness ships natively. Autonomy enforcement, quality gates, parallel subagent dispatch, and a session resume protocol sit on top of whatever the harness provides.
Instead of just telling an assistant about your codebase, Spellbook gives the assistant structured workflows for research, design, implementation, testing, and review, along with guardrails for the specific ways LLMs tend to cut corners.
The orchestrator pattern
The main agent dispatches subagents rather than doing implementation work directly. This keeps the main context window free for strategic coordination instead of filling it with source code, and it means each subagent starts with a fresh perspective rather than carrying accumulated assumptions. Parallel dispatch lets multiple tasks run simultaneously.
Epistemic rigor
The system is designed to distrust its own outputs. [Fact-checking][fact-checking] treats every claim as a hypothesis to verify. [Green mirage auditing][auditing-green-mirage] asks whether a test would actually fail if the code were broken, which is a different question from whether the test passes. [Hunch verification][verifying-hunches] intercepts moments of claimed discovery and requires reframing them as testable hypotheses. [Dehallucination][dehallucination] names the specific ways LLMs confabulate and provides recovery protocols.
Test-driven development is treated as an epistemic practice: tests written before implementation answer "what should this do?" while tests written after answer "what does this do?" -- and the system's quality gates are designed around that difference.
Hallucination prevention draws on peer-reviewed research. [Chain-of-Verification][cove-protocol] self-interrogation (Dhuliawala et al., 2023) requires verification skills to generate and answer questions about their own claims before finalizing verdicts. [Atomic claim decomposition][decompose-claims] (Min et al., FActScore, EMNLP 2023) breaks compound statements into independently verifiable units. API hallucination detection checklists in [code review][code-review] and [quality enforcement][enforcing-code-quality] catch the specific pattern where LLMs generate syntactically valid but non-existent API calls.
Named failure modes
LLMs fail in predictable ways, and Spellbook names those patterns so it can build mechanical countermeasures. Seven rationalization patterns are catalogued and blocked. Three consecutive fix failures trigger architectural reassessment instead of a fourth attempt. Research stagnation triggers a plateau breaker. A devil's advocate review that finds zero issues is flagged as incomplete.
Quality gates
Every substantial skill runs as a sequence of phases with mandatory gates between them. Tests must pass, code review must clear, claims must verify against source, and tests must actually catch regressions. These gates cannot be bypassed by YOLO mode or autonomy settings -- YOLO grants permission to act without asking, but not to skip verification.
Composition
Skills invoke skills. [develop][develop] orchestrates [design-exploration], [writing-plans], [test-driven-development], [requesting-code-review], [fact-checking], [auditing-green-mirage], and [finishing-a-development-branch]. [debugging][debugging] invokes [verifying-hunches] and [isolated-testing]. When a skill outgrows its scope, it splits into a thin orchestrator and supporting commands.
Self-improvement
Some skills exist to improve other skills. [Usage analytics][analyzing-skill-usage] measure completion and correction rates. The [skill-writing skill][writing-skills] applies TDD to skill creation itself. [Instruction engineering][instruction-engineering] codifies prompt research into technique, and [prompt sharpening][sharpening-prompts] audits for ambiguity. A/B testing compares skill versions against each other so improvements can be measured rather than assumed.
The develop Skill
You say "add dark mode" or "migrate the auth system to OAuth2" or "build a webhook delivery pipeline with retry logic." The [develop][develop] skill orchestrates the full feature lifecycle through 20+ specialized skills and commands. The first question it asks is how involved you want to be:
Fully autonomous. Describe the feature and walk away. It researches your codebase, surfaces ambiguities, resolves them, designs the architecture, writes a detailed implementation plan, builds with test-driven development, reviews its own code, fact-checks its claims, audits its tests for false confidence, and opens a PR. Each step runs in a fresh subagent with its own quality gate.
Highly interactive. The same pipeline with the same rigor, but you stay in the conversation. Ambiguities become specific questions grounded in what it found in your code, and architectural tradeoffs come with evidence. Checkpoints pause for your input.
Or anywhere between. You can run mostly autonomous with pauses only for critical decisions, and set the level once at the start.
How it works
The system classifies your request by complexity using mechanical heuristics -- file count, behavioral change, test impact, structural change, integration points. Trivial changes exit the skill entirely. Simple changes follow a lightwei
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
86.1kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.1kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
ruflo
73.5k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
CowAgent
47.2kOpen-source super AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
