cache-efficiency
Analyze prompt-cache effectiveness for Claude Code usage from the Agent Monitor dashboard — cache hit rate (total_cache_read / (total_cache_read + total_input)), cache_write vs cache_read reuse, cache-read vs cache-write spend, and the sessions with the poorest reuse.
Install / Use
npx skills add hoangsonww/Claude-Code-Agent-Monitor --skill cache-efficiencyInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
OperationsSupported Platforms
Our assessment of cache-efficiency
cache-efficiency scores 85/100 on our quality scale, 479th of 736 Operations skills we index.
Its SKILL.md is 3.8 KB long, well organised into 11 sections with 1 code example: a solid amount of guidance for an agent.
With 1,015 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 12 days ago, so cache-efficiency is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
cache-efficiency compared with similar skills
All 4 of these similar skills score higher than cache-efficiency; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| cache-efficiency (this skill)by hoangsonww | 85 | 1.0k | 12d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 92.4k | 21d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.5k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.9k | today | MCP Server |
| crawl4aiby unclecode | 100 | 84.8k | 1d ago | MCP Server |
Frequently asked questions
- How do I install cache-efficiency?
- Run
npx skills add hoangsonww/Claude-Code-Agent-Monitor --skill cache-efficiency. The install tabs above show the steps for each supported agent. - Which AI agents does cache-efficiency work with?
- It is written for Claude Code, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is cache-efficiency safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is cache-efficiency still maintained?
- The repository was last updated 12 days ago, so cache-efficiency is actively maintained.
Skill content
View source on GitHubname: cache-efficiency description: > Analyze prompt-cache effectiveness for Claude Code usage from the Agent Monitor dashboard — cache hit rate (total_cache_read / (total_cache_read + total_input)), cache_write vs cache_read reuse, cache-read vs cache-write spend, and the sessions with the poorest reuse. Pulls token totals from /api/analytics, per-session detail from /api/sessions, and dollar splits from /api/pricing/cost. Use when diagnosing cache spend or deciding whether prompt caching is paying off.
Cache Efficiency
Diagnose whether prompt caching is actually saving money, and where it is not.
Input
The user provides: $ARGUMENTS
This may be: empty (analyze the whole fleet), "today" / "this week" / a date range, a session ID to scope the analysis, or a target like "hit rate > 80%". When empty, analyze all data from /api/analytics.
Data Sources
| Endpoint | Returns |
|----------|---------|
| GET /api/analytics | tokens.total_input, tokens.total_output, tokens.total_cache_read, tokens.total_cache_write (baselines pre-summed), plus daily_sessions |
| GET /api/sessions?limit=200 | Session list — each has model, cwd, started_at, ended_at, inline cost, metadata (JSON: usage_extras with cache token detail) |
| GET /api/sessions/{id} | Full session detail with nested agents and events, for drill-down on a flagged session |
| GET /api/pricing/cost | { total_cost, breakdown: [{ model, input_tokens, output_tokens, cache_read_tokens, cache_write_tokens, cost, matched_rule }] } — used to price cache read vs write spend |
How cache economics work
cache_hit_rate = total_cache_read / (total_cache_read + total_input)
cache_reuse = total_cache_read / total_cache_write
cache_read_cost = (cache_read_tokens / 1M) × cache_read_per_mtok
cache_write_cost = (cache_write_tokens / 1M) × cache_write_per_mtok
Cache writes cost more per token than cache reads (e.g. Sonnet $3.75 write vs $0.30 read per Mtok), and writes are billed even if the cached block is never reused. The payoff only arrives on subsequent reads — so a healthy fleet shows cache_read_tokens far exceeding cache_write_tokens. When cache_reuse < 1, you are paying to cache context you barely re-read.
Token counts are effective totals = current + baseline (baselines preserve pre-compaction tokens).
Report Sections
1. Fleet Cache Hit Rate
From /api/analytics: compute cache_hit_rate × 100. State raw total_cache_read and total_input. Benchmark: >70% strong, 40–70% moderate, <40% weak prompt-cache utilization.
2. Write vs Read Reuse
Compute cache_reuse = total_cache_read / total_cache_write. Show both token counts. Flag if reuse < 1 (writing more cache than is ever read back).
3. Cache Spend Split
From /api/pricing/cost breakdown, sum cache_read_cost and cache_write_cost across all models. Show the dollar split and what fraction of total cost is cache-write overhead vs cache-read savings.
4. Sessions With Poor Reuse
From /api/sessions?limit=200, parse metadata.usage_extras for per-session cache read/write where available; rank sessions by lowest read/write reuse (and by cache_write-heavy cost). List the worst 10 with model, cost, and reuse ratio. Use /api/sessions/{id} to drill into any single flagged session.
5. Recommendations
- Sessions where
cache_write >> cache_read: short or one-shot sessions rarely recoup cache writes — note them. - Stable, repeated context (system prompts, large files) should be cached once and reused; high churn defeats caching.
- Estimate the dollar impact of raising the hit rate to the next benchmark tier.
Output
Structured Markdown with tables. Currency as USD to 4 decimal places; rates as $/Mtok; percentages with ▲/▼ for any trend. Token counts with thousands separators.
Related Skills
Agent-Reach
92.4kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.5kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Scrapling
85.9k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
crawl4ai
84.8kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
