caveman
๐ชจ why use many token when few token do trick. Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
Install / Use
npx skills add JuliusBrussee/cavemanInstalls into whichever agent you are using.
CLAUDE.md
Claude Code project instructions
Quality Score
Category
AI & Machine LearningSupported Platforms
Tags
Skill content
View source on GitHubCaveman is a skill/plugin for Claude Code, Codex, Gemini, Cursor, Windsurf, Cline, Copilot, and 30+ other agents. Install once. Agent drops the filler and answers in tight caveman-speak, keeping code, commands, and errors byte-for-byte exact. You save output tokens on every reply, forever.
Before / After
<table> <tr> <th width="50%">๐ฃ๏ธ Normal agent โ 69 tokens</th> <th width="50%"><img src="docs/assets/dancing-rock.svg" width="18" height="18" alt=""> Caveman agent โ 19 tokens</th> </tr> <tr> <td valign="top"></td> <td valign="top">The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I'd recommend using useMemo to memoize the object.
</td> </tr> <tr> <td valign="top">New object ref each render. Inline object prop = new ref = re-render. Wrap in
useMemo.
</td> <td valign="top">Sure! I'd be happy to help you with that. The issue you're experiencing is most likely caused by your authentication middleware not properly validating the token expiry. Let me take a look and suggest a fix.
</td> </tr> </table>Bug in auth middleware. Token expiry check use
<not<=. Fix:
Same fix. Third of the words. Nothing technical lost.
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ output tokens saved โโโโโโโโโ 65% โ
โ input tokens saved โโโโโโโโโ 0% โ
โ technical accuracy โโโโโโโโโ 100% โ
โ vibes โโโโโโโโโ OOG โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Caveman no make brain smaller. Caveman make mouth smaller. Shrinks what the agent says, not what it knows.
That 65% is the prose number, measured on replies like the ones above. On a full agentic coding run, where most of the output is code and tool calls, it's 8.5%. Same skill, different workload โ mechanism below.
Install
One command. Finds every agent on your machine. Installs for each.
# macOS ยท Linux ยท WSL ยท Git Bash
curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash
# Windows ยท PowerShell 5.1+
irm https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.ps1 | iex
~30 seconds. Needs Node โฅ18. Skips agents you no have. Safe to re-run.
Prefer one agent at a time? Each has its own path:
# Claude Code plugin
claude plugin marketplace add JuliusBrussee/caveman && claude plugin install caveman@caveman
# Gemini CLI extension
gemini extensions install https://github.com/JuliusBrussee/caveman --consent
# Cursor / Windsurf / Cline / Codex / 30+ more, via the skills registry
npx skills add JuliusBrussee/caveman -a cursor
The full per-agent matrix, all flags, dry-run, and uninstall live in INSTALL.md.
[!TIP] Turn it on: type
/cavemanor say "talk like caveman". Turn it off: say "normal mode". On Claude Code, Codex, and Gemini it's already on from message one. No command needed.
Install broke? Open your agent in this repo and say: "Read CLAUDE.md and INSTALL.md, install caveman for me." Agent read repo, agent fix own brain. Snake eat tail.
Pick your grunt
Six levels. Switch anytime with /caveman <level>. Level sticks until you change it or the session ends.
| Level | Same sentence, shrunk |
|---|---|
| normal agent | You should wrap the object in useMemo, since a new reference is created on every render. |
| lite | Wrap object in useMemo. New ref created every render. |
| full (default) | New ref each render. Wrap object in useMemo. |
| ultra | New ref/render. useMemo it. |
| wenyan | New ref every render, so wrap in useMemo โ rendered in classical Chinese, shorter still. |
[!NOTE] Speak your tongue. Caveman keeps your language. Write Portuguese, caveman grunt Portuguese. Spanish, French, same. It compresses the style, never translates.
wenyanmode is the exception on purpose: classical Chinese packs the most meaning per token.
What you get
| Command | What it does |
|---|---|
| /caveman [lite\|full\|ultra\|wenyan] | Compress every reply. Level sticks for the session. |
| /caveman-commit | Conventional Commit messages, โค50-char subject. Why over what. |
| /caveman-review | One-line PR comments: L42: ๐ด bug: user null. Add guard. |
| /caveman-stats | Real session token usage, lifetime savings, USD. Tweetable line with --share. |
| /caveman-compress <file> | Rewrite a memory file (like CLAUDE.md) into caveman-speak. Cuts ~46% input tokens every session after. Code, URLs, paths byte-preserved. |
| caveman-shrink | MCP middleware. Wraps any MCP server, compresses its tool descriptions. npm. |
| cavecrew-* | Caveman subagents (investigator, builder, reviewer). ~60% fewer tokens than vanilla, so main context lasts longer. |
[!TIP] On Claude Code the statusline shows
[CAVEMAN] โ 12.4kโ that's your lifetime tokens saved, updated on every/caveman-stats. Silence it withCAVEMAN_STATUSLINE_SAVINGS=0.
Benchmarks
Real token counts from the Claude API. Average 65% output reduction across 10 chat-style prompts (range 22โ87%), measured against default verbose replies. Output tokens only, committed and reproducible in benchmarks/ and evals/. This is one-question-one-answer, not a full agentic coding run โ for that number, see JetBrains below.
| Task | Normal | Caveman | Saved | |------|-------:|--------:|------:| | Explain React re-render bug | 1180 | 159 | 87% | | Fix auth middleware token expiry | 704 | 121 | 83% | | Set up PostgreSQL connection pool | 2347 | 380 | 84% | | Explain git rebase vs merge | 702 | 292 | 58% | | Refactor callback to async/await | 387 | 301 | 22% | | Architecture: microservices vs monolith | 446 | 310 | 30% | | Review PR for security issues | 678 | 398 | 41% | | Docker multi-stage build | 1042 | 290 | 72% | | Debug PostgreSQL race condition | 1200 | 232 | 81% | | Implement React error boundary | 3454 | 456 | 87% | | Average | 1214 | 294 | 65% |
<!-- BENCHMARK-TABLE-END -->[!IMPORTANT] Honest number warning. Caveman only shrinks output tokens. Input and reasoning tokens are untouched, and the skill itself adds ~1โ1.5k input tokens per turn. So whole-session savings run smaller than the output number, and on already-terse workloads they can go net-negative. The real win is readability and speed. Cost savings are the bonus. When caveman wins, when it loses, and how to measure it yourself: docs/HONEST-NUMBERS.md.
Independently measured: JetBrains, 86 tasks
JetBrains ran the skill against 86 tasks from SkillsBench in July 2026 โ real coding work, auto-graded by each task's own tests, Claude Code on claude-sonnet-5, skill forced on for every reply.
| Workload | Output tokens saved | Measured by | |---|---:|---| | Chat-style prose | 65% | us, table above | | Agentic coding run | 8.5% | JetBrains, 86 tasks |
Both numbers are real. They measure different workloads, and the gap is mechanical: caveman compresses narration and leaves code, diffs, tool calls, and error strings byte-exact. In a chat answer, narration is the whole reply. In an agentic run it's the thin layer between tool calls, so that's all there is to squeeze. An output-only skill has a low ceiling on work that is mostly not prose.
Pick the number that matches your workload:
- Agent writes you prose โ explanations, review, docs, debugging walkthroughs โ 65% territory.
- Agent works a repo unattended โ single digits. Not zero, not 65%.
Quality was unaffected: across 86 auto-graded tasks the two arms were statistically indistinguishable. Small mouth, same brain โ checked by someone who didn't ship it.
Two things follow:
- Agentic bills are mostly input tokens, which an output-only skill cannot touch by construction.
/caveman-compressandcaveman-shrinkchip at that side; the skill alone never will. - The right number is your number. JetBrains had to run a full paid benchmark to find out what caveman does on their stack. That's the job Caveman 2 exists to do โ for yours, continuously.
Turns out short isn't just cheaper. A March 2026 paper, Brevity Constraints Reverse Performance Hierarchies in Language Models, tested 31 models and found that constraining large models to brief answers improved accuracy by ~26 points on some benchmarks. Sometimes less word = more correct.
<details> <summary><strong>caveman-compress receipts</strong> โ real memory files, cutting input tokens forever</summary> <br>| File | Original | Compressed | Saved |
|---|---:|---:|---:|
| claude-md-preferences.md | 706 | 285 | 59.6% |
| project-notes.md | 1145 | 535 | 53.3% |
| claude-md-project.md | 1122 | 636 | 43.3% |
| todo-list.md | 627 | 388 | 38.1% |
| mixed-with-code.md | 888 | 560 | 36.9% |
| Average | 898 | 481 | 46% |
Every session after, that file loads ~46% smaller. Input tokens saved forever, not just one reply.
</details>The whole cave
<table> <tr><td><img src="docs/assets/dancing-rock.svg" width="20" height="20" alt=""> Want the whole agent, not just its mouth? โ caveman-code
This skill shrinks what an agent says. caveman-code shrinks everything โ a full terminal coding agent, caveman top to bottom. ~2ร fewer tokens than Codex on identical tasks. 20+ providers, plan mode, autopilot goal loop, MIT.
npm install -g @juliusbrussee/caveman-code
</td></tr>
</table>
Five tools, one idea: **agent do mor
Truncated for display โ read the full file on GitHub.
Related Skills
claude-mem
94.4kPersistent Context Across Sessions for Every Agent โ Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
84.2kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu โ one CLI, zero API fees.
Understand-Anything
83.5kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
73.4kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
