SkillAgentSearch skills...

regression-watch

Detect quality and efficiency regressions over time using Agent Monitor data — rising error rate (APIError events), falling cache hit rate, growing compaction frequency, and climbing cost-per-session.

Install / Use

npx skills add hoangsonww/Claude-Code-Agent-Monitor --skill regression-watch

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

81/100

Category

Operations

Supported Platforms

Claude Code

Our assessment of regression-watch

regression-watch scores 81/100 on our quality scale, 619th of 751 Operations skills we index.

Its SKILL.md is 4.0 KB long, well organised into 11 sections and no code examples: a solid amount of guidance for an agent.

With 1,015 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
26/30
Structure
13/20
Description
15/15
Adoption
13/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 10 days ago, so regression-watch is actively maintained.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

regression-watch compared with similar skills

All 4 of these similar skills score higher than regression-watch; compare them before choosing.

SkillScoreStarsUpdatedFormat
regression-watch (this skill)by hoangsonww811.0k10d agoSKILL.md
Agent-Reachby Panniantong10089.8k18d agoCLAUDE.md
headroomby headroomlabs-ai10074.4ktodayCLAUDE.md
Scraplingby D4Vinci10085.5ktodayMCP Server
crawl4aiby unclecode10084.7k8d agoMCP Server

Frequently asked questions

How do I install regression-watch?
Run npx skills add hoangsonww/Claude-Code-Agent-Monitor --skill regression-watch. The install tabs above show the steps for each supported agent.
Which AI agents does regression-watch work with?
It is written for Claude Code, as a SKILL.md file. Other agents that read the same format can often use it too.
Is regression-watch safe to use?
It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is regression-watch still maintained?
The repository was last updated 10 days ago, so regression-watch is actively maintained.

name: regression-watch description: > Detect quality and efficiency regressions over time using Agent Monitor data — rising error rate (APIError events), falling cache hit rate, growing compaction frequency, and climbing cost-per-session. Splits history into an earlier baseline window and a recent window and reports which metrics are getting worse, by how much, and where. Use when checking whether things are degrading or trending in the wrong direction.

Regression Watch

Detect whether Claude Code sessions are getting worse over time across quality and efficiency metrics, using Agent Monitor data.

Input

The user provides: $ARGUMENTS

This may be:

  • empty or "all" — check every regression metric (default)
  • "errors" — error-rate regression only
  • "cache" — cache hit-rate regression only
  • "compaction" — compaction-frequency regression only
  • "cost" — cost-per-session regression only
  • A window like "last 30d" or "30 vs 90" — set the recent vs baseline window sizes

Data Sources

| Endpoint | Returns | |----------|---------| | GET /api/analytics | daily_events (365d), daily_sessions (365d), event_types, tokens (total_input, total_output, total_cache_read, total_cache_write — baselines pre-summed), avg_events_per_session | | GET /api/events?session_id=X | Event stream incl. APIError, Compaction, PreToolUse/PostToolUse — used to localize regressions to specific sessions | | GET /api/pricing/cost | { total_cost, breakdown[...] } — total cost to derive cost-per-session | | GET /api/pricing/cost/{sessionId} | Per-session cost — used to compare recent vs baseline session cost | | GET /api/workflows/{sessionId} | compaction (impact), errorPropagation (by depth), effectiveness — per-session quality signals | | GET /api/sessions?limit=N | Sessions with started_at, cost, metadata — to bucket sessions into time windows |

Report Sections

1. Windowing

Split history into a baseline window (older) and a recent window (newer). Default: recent = last 30 days, baseline = the 30–90 day range before it. Use daily_events/daily_sessions for series metrics and GET /api/sessions?limit=N to assign sessions to each window by started_at.

2. Error Rate Regression

  • Recent error rate = APIError count / total events in the recent window (from event_types and daily_events, or per-session GET /api/events).
  • Compare to the baseline rate. Flag if recent is higher.
  • Report the absolute and relative change and which sessions contributed most APIError events.

3. Cache Hit Rate Regression

  • Cache hit rate = total_cache_read / (total_cache_read + total_input).
  • Compute for each window (per-window input/cache_read from session metadata or the pricing breakdown). Flag a falling hit rate — that means more uncached input tokens and higher cost.

4. Compaction Frequency Regression

  • Compaction frequency = Compaction events / session per window (from event_types / daily_events, confirmed via per-session GET /api/workflows/{id} compaction). Flag a rising rate — context is overflowing more often.

5. Cost-per-Session Regression

  • Cost-per-session = window total cost / window session count, using GET /api/pricing/cost overall and GET /api/pricing/cost/{id} for the sessions in each window. Flag a climbing value.

6. Verdict

Roll up which metrics regressed, rank by relative worsening, and name the most likely driver (e.g., cache hit rate fell → cost per session climbed).

Output

  • A Markdown table: metric | baseline | recent | Δ | direction (▲ worse / ▼ better) | verdict.
  • Tag each regressed metric 🔴 (clear regression), 🟡 (mild/within noise), or 🟢 (improved).
  • Currency in USD to 4 decimals; rates as percentages to 2 decimals.
  • List the specific session IDs that contributed most to any regression.
  • End with the single highest-priority regression to address and a concrete next step.
  • Read-only: only report what the API returns; never fabricate baselines.

Related Skills

View on GitHub
GitHub Stars1.0k
CategoryOperations
Updated10d ago
Forks238

Languages

JavaScript

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions