regression-watch
Detect quality and efficiency regressions over time using Agent Monitor data — rising error rate (APIError events), falling cache hit rate, growing compaction frequency, and climbing cost-per-session.
Install / Use
npx skills add hoangsonww/Claude-Code-Agent-Monitor --skill regression-watchInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
OperationsSupported Platforms
Our assessment of regression-watch
regression-watch scores 81/100 on our quality scale, 619th of 751 Operations skills we index.
Its SKILL.md is 4.0 KB long, well organised into 11 sections and no code examples: a solid amount of guidance for an agent.
With 1,015 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 10 days ago, so regression-watch is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
regression-watch compared with similar skills
All 4 of these similar skills score higher than regression-watch; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| regression-watch (this skill)by hoangsonww | 81 | 1.0k | 10d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 89.8k | 18d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.4k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.5k | today | MCP Server |
| crawl4aiby unclecode | 100 | 84.7k | 8d ago | MCP Server |
Frequently asked questions
- How do I install regression-watch?
- Run
npx skills add hoangsonww/Claude-Code-Agent-Monitor --skill regression-watch. The install tabs above show the steps for each supported agent. - Which AI agents does regression-watch work with?
- It is written for Claude Code, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is regression-watch safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is regression-watch still maintained?
- The repository was last updated 10 days ago, so regression-watch is actively maintained.
Skill content
View source on GitHubname: regression-watch description: > Detect quality and efficiency regressions over time using Agent Monitor data — rising error rate (APIError events), falling cache hit rate, growing compaction frequency, and climbing cost-per-session. Splits history into an earlier baseline window and a recent window and reports which metrics are getting worse, by how much, and where. Use when checking whether things are degrading or trending in the wrong direction.
Regression Watch
Detect whether Claude Code sessions are getting worse over time across quality and efficiency metrics, using Agent Monitor data.
Input
The user provides: $ARGUMENTS
This may be:
- empty or "all" — check every regression metric (default)
- "errors" — error-rate regression only
- "cache" — cache hit-rate regression only
- "compaction" — compaction-frequency regression only
- "cost" — cost-per-session regression only
- A window like "last 30d" or "30 vs 90" — set the recent vs baseline window sizes
Data Sources
| Endpoint | Returns |
|----------|---------|
| GET /api/analytics | daily_events (365d), daily_sessions (365d), event_types, tokens (total_input, total_output, total_cache_read, total_cache_write — baselines pre-summed), avg_events_per_session |
| GET /api/events?session_id=X | Event stream incl. APIError, Compaction, PreToolUse/PostToolUse — used to localize regressions to specific sessions |
| GET /api/pricing/cost | { total_cost, breakdown[...] } — total cost to derive cost-per-session |
| GET /api/pricing/cost/{sessionId} | Per-session cost — used to compare recent vs baseline session cost |
| GET /api/workflows/{sessionId} | compaction (impact), errorPropagation (by depth), effectiveness — per-session quality signals |
| GET /api/sessions?limit=N | Sessions with started_at, cost, metadata — to bucket sessions into time windows |
Report Sections
1. Windowing
Split history into a baseline window (older) and a recent window (newer).
Default: recent = last 30 days, baseline = the 30–90 day range before it. Use
daily_events/daily_sessions for series metrics and GET /api/sessions?limit=N
to assign sessions to each window by started_at.
2. Error Rate Regression
- Recent error rate =
APIError count / total eventsin the recent window (fromevent_typesanddaily_events, or per-sessionGET /api/events). - Compare to the baseline rate. Flag if recent is higher.
- Report the absolute and relative change and which sessions contributed most
APIErrorevents.
3. Cache Hit Rate Regression
- Cache hit rate =
total_cache_read / (total_cache_read + total_input). - Compute for each window (per-window input/cache_read from session metadata or the pricing breakdown). Flag a falling hit rate — that means more uncached input tokens and higher cost.
4. Compaction Frequency Regression
- Compaction frequency =
Compaction events / sessionper window (fromevent_types/daily_events, confirmed via per-sessionGET /api/workflows/{id}compaction). Flag a rising rate — context is overflowing more often.
5. Cost-per-Session Regression
- Cost-per-session = window total cost / window session count, using
GET /api/pricing/costoverall andGET /api/pricing/cost/{id}for the sessions in each window. Flag a climbing value.
6. Verdict
Roll up which metrics regressed, rank by relative worsening, and name the most likely driver (e.g., cache hit rate fell → cost per session climbed).
Output
- A Markdown table: metric | baseline | recent | Δ | direction (▲ worse / ▼ better) | verdict.
- Tag each regressed metric 🔴 (clear regression), 🟡 (mild/within noise), or 🟢 (improved).
- Currency in USD to 4 decimals; rates as percentages to 2 decimals.
- List the specific session IDs that contributed most to any regression.
- End with the single highest-priority regression to address and a concrete next step.
- Read-only: only report what the API returns; never fabricate baselines.
Related Skills
Agent-Reach
89.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.4kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Scrapling
85.5k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
crawl4ai
84.7kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
