SkillAgentSearch skills...

benchmark

Benchmark one session (or a small recent set) against the rolling average using Agent Monitor data — cost, total tokens, tool count, and workflow complexity score — and report where each metric lands as a percentile of the population. Tells you whether a session was normal, cheap, or an outlier

Install / Use

npx skills add hoangsonww/Claude-Code-Agent-Monitor --skill benchmark

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

81/100

Category

Automation

Supported Platforms

Universal

Tags

Our assessment of benchmark

benchmark scores 81/100 on our quality scale, 2346th of 2,843 Automation skills we index.

Its SKILL.md is 3.3 KB long, well organised into 9 sections and no code examples: a solid amount of guidance for an agent.

With 1,015 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
26/30
Structure
13/20
Description
15/15
Adoption
13/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 10 days ago, so benchmark is actively maintained.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

benchmark compared with similar skills

All 4 of these similar skills score higher than benchmark; compare them before choosing.

SkillScoreStarsUpdatedFormat
benchmark (this skill)by hoangsonww811.0k10d agoSKILL.md
Agent-Reachby Panniantong10089.8k18d agoCLAUDE.md
Scraplingby D4Vinci10085.5ktodayMCP Server
rufloby ruvnet10073.8ktodayMCP Server
algorithmic-artby anthropics100177.9k11d agoSKILL.md

Frequently asked questions

How do I install benchmark?
Run npx skills add hoangsonww/Claude-Code-Agent-Monitor --skill benchmark. The install tabs above show the steps for each supported agent.
Which AI agents does benchmark work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is benchmark safe to use?
It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is benchmark still maintained?
The repository was last updated 10 days ago, so benchmark is actively maintained.

name: benchmark description: > Benchmark one session (or a small recent set) against the rolling average using Agent Monitor data — cost, total tokens, tool count, and workflow complexity score — and report where each metric lands as a percentile of the population. Tells you whether a session was normal, cheap, or an outlier. Use when judging whether a session was typical or out of band.

Benchmark

Score a session against the rolling population average and report its percentile on cost, tokens, tool count, and complexity using Agent Monitor data.

Input

The user provides: $ARGUMENTS

This may be:

  • A single session ID — benchmark that session
  • "latest" — benchmark the most recent session
  • "latest N" — benchmark the N most recent sessions, each vs the average
  • empty — benchmark the most recent session (default)

Data Sources

| Endpoint | Returns | |----------|---------| | GET /api/sessions?limit=N | Population of sessions with cost, model, started_at, metadata (turn_count, total_turn_duration_ms) — builds the rolling baseline | | GET /api/pricing/cost/{sessionId} | { total_cost, breakdown:[{ input_tokens, output_tokens, cache_read_tokens, cache_write_tokens, cost }] } — the target session's cost and tokens | | GET /api/workflows/{sessionId} | complexity (score), stats (tool/event counts), toolFlow (distinct tools used) — the target session's tool count and complexity | | GET /api/analytics | avg_events_per_session, tool_usage, daily_sessions — corroborates population-level averages |

Report Sections

1. Build the Baseline

Fetch the population with GET /api/sessions?limit=200 (the rolling set). For each session gather cost (GET /api/pricing/cost/{id} or the list cost field), total tokens (sum of the 4 token types from the pricing breakdown), tool count and complexity (GET /api/workflows/{id}). Compute mean, median, and standard deviation for each metric across the population.

2. Measure the Target

For the requested session, pull the same four metrics:

  • Cost — total_cost from GET /api/pricing/cost/{id}.
  • Total tokens — input + output + cache_read + cache_write summed from the breakdown.
  • Tool count — distinct/total tools from GET /api/workflows/{id} stats/toolFlow.
  • Complexity score — complexity.score from GET /api/workflows/{id}.

3. Percentile and Deviation

For each metric report the target's percentile within the population (share of sessions at or below it) and its z-score (value − mean) / stddev. Label each: below average / typical / above average / outlier (|z| > 2).

4. Verdict

State whether the session was normal overall. If it is an outlier, name which metric drove it (e.g., complexity p96, cost p91 → an unusually heavy session).

Output

  • A Markdown table: metric | session value | population mean | percentile | z-score | label.
  • Currency in USD to 4 decimals; tokens and tool counts as integers; complexity to 2 decimals.
  • Use ▲ for above-average and ▼ for below-average vs the mean.
  • One-line verdict: "Normal session" or "Outlier — driven by <metric> (pNN)".
  • When benchmarking multiple sessions, one row block per session plus a summary line.
  • Read-only: percentiles come only from the fetched population; never fabricate the baseline.

Related Skills

View on GitHub
GitHub Stars1.0k
CategoryAutomation
Updated10d ago
Forks238

Languages

JavaScript

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions