SkillAgentSearch skills...

Agent Cost Report

Believable agent cost report for any period, default the last 7 full days PT, not counting today. Measured tokens from Claude Code transcripts priced at OpenRouter list prices (ESTIMATED), measured provider spend when a sanctioned source exists, note-taker cost separate, Timing-style HTML/PDF plus r…

Install / Use

npx skills add thedotmack/claude-mem --skill agent-cost-report

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

97/100

Supported Platforms

Claude Code

Our assessment of Agent Cost Report

Agent Cost Report scores 97/100 on our quality scale, 25th of 597 Data & Analytics skills we index (top 5%).

Its SKILL.md is 20 KB long, well organised into 21 sections with 1 code example: a thorough specification that gives an agent plenty to work with.

With 97,136 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
17/20
Description
15/15
Adoption
20/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated yesterday, so Agent Cost Report is actively maintained.
  • It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Agent Cost Report compared with similar skills

All 4 of these similar skills score higher than Agent Cost Report; compare them before choosing.

SkillScoreStarsUpdatedFormat
Agent Cost Report (this skill)by thedotmack9797.1k1d agoSKILL.md
Agent-Reachby Panniantong10093.0k22d agoCLAUDE.md
headroomby headroomlabs-ai10074.6ktodayCLAUDE.md
CowAgentby zhayujie10047.3ktodayCLAUDE.md
Scraplingby D4Vinci10086.1ktodayMCP Server

Frequently asked questions

How do I install Agent Cost Report?
Run npx skills add thedotmack/claude-mem --skill "Agent Cost Report". The install tabs above show the steps for each supported agent.
Which AI agents does Agent Cost Report work with?
It is written for Claude Code, as a SKILL.md file. Other agents that read the same format can often use it too.
Is Agent Cost Report safe to use?
It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is Agent Cost Report still maintained?
The repository was last updated yesterday, so Agent Cost Report is actively maintained.

name: Agent Cost Report description: >- Believable agent cost report for any period, default the last 7 full days PT, not counting today. Measured tokens from Claude Code transcripts priced at OpenRouter list prices (ESTIMATED), measured provider spend when a sanctioned source exists, note-taker cost separate, Timing-style HTML/PDF plus report.json, line-items.csv, evidence.json. allowed-tools:

  • Bash
  • Read
  • Write
  • AskUserQuestion
  • mcp__plugin_claude-mem_mcp-search__search
  • mcp__plugin_claude-mem_mcp-search__timeline
  • mcp__plugin_claude-mem_mcp-search__get_observations

Agent Cost Report

Claude-Mem / Claude Code skill. Runtime is the scripts/ pipeline (transcripts → tokens → dollars → Timing-style report) plus a progressive Mem Search review pass that confirms the drafted labels. The Notion draft is SPEC history only — never the product, never the runtime, never the ship vehicle.

Resolve the absolute directory containing this SKILL.md; all helper paths are relative to that directory. ${CLAUDE_SKILL_DIR} is the shortcut: python3 "${CLAUDE_SKILL_DIR}/scripts/acr.py" …. Python 3.9+ standard library only, with IANA timezone data for America/Los_Angeles; no pip installs. tzdata is the only exception to the no-pip-installs rule, and only when acr.py exits saying that timezone data is unavailable (stock Windows Python ships none): show the user the lines it printed and ask (AskUserQuestion) before installing it. On a yes, run the line for your shell (the plain quoted line in Bash or cmd, the PowerShell: line in PowerShell), then rerun the failed command. Setting PYTHONTZPATH instead needs no install. The look lives in scripts/acr/render.py, never here.

Purpose

Turn Claude-Mem activity into a manager-readable cost and failure report.

Product idea: a reusable skill that searches Claude-Mem via Mem Search, reconstructs real units of work, assigns cost and failure categories, and renders a printable report.

Insight north star

The headline is dollars, to two decimals, labeled. The dollars come from Claude Code transcripts (exact per-reply token usage) priced at OpenRouter public list prices, so they are ESTIMATED. Measured provider spend appears only when a sanctioned source gives it. Directly under the dollars: what the mistakes cost, what shipped, and both on one time axis.

Questions the report must answer

  1. What work was completed?
  2. What did each outcome cost?
  3. What was wasted through looping, hedging, wrong turns, rework, or poor routing?
  4. Were any unauthorized actions attempted?
  5. What should the manager change next?

Primary unit = cost per completed outcome (not cost per observation).

When to use

  • "Agent cost report" / "cost per outcome" / "failure economics" / "was this session worth it" / "what did the agents cost this week"
  • After a real Mem session dig when leadership needs outcome economics
  • Sample / ship packs that need self-contained HTML + JSON + CSV + evidence

Memory dig mechanics: the claude-mem mem-search skill (progressive recall).

Default scope (ALWAYS)

  • Unless the user names a specific session / range / project, the window is the last 7 full days in PT, not counting today: end = the PT midnight that started today (exclusive), start = end − 7 days. The default window never contains a partial day (G3, Alex 2026-09-25).
  • Explicit windows: --start YYYY-MM-DD --end YYYY-MM-DD (PT calendar days, end exclusive). An explicit --end later than today marks the last day "partial, generated HH:MM PT".
  • One session: --session <content_session_id>. One project plus a period: --project <name> --start … --end … (worktrees of the project are included).
  • The report Scope strip shows the PT range, and Details list every session id in scope.

Progressive Mem Search (ALWAYS)

Follow the claude-mem mem-search three layers. Keep spend light.

  1. Search — get an index of IDs (titles, types, token hints).
  2. Timeline — only around anchors you care about.
  3. Observations — get_observations for the filtered IDs you will cite as evidence.

Recipe:

  1. Resolve scope (default: the last 7 full PT days).
  2. Search → collect IDs.
  3. Timeline for thin context only.
  4. Observations for intended / actual / outcome / waste / rework / blocked / unauthorized / status.
  5. Group into named work items + failure events.
  6. Calculate line-item costs.
  7. Render HTML + optional PDF (+ json/csv/evidence).
  8. Keep evidence IDs in the appendix — do not dump entire timelines into the main report.

ADHD process bullets:

  • Search first → pick IDs → timeline only if context is thin → fetch only needed obs.
  • Work-item titles are invented for managers ("Restore search after Chroma crash-loop"); observation titles stay evidence-only.
  • Evidence appendix lists obs IDs + short titles; main sections stay outcome-first.

In this skill the search pass is the review step (see Recipe step 4): the pipeline drafts categories and failure types from keywords (label_source: keyword); the orchestrator confirms or changes each line item's category and failure_type from its cited evidence IDs and applies the result with review --apply. Items left unreviewed keep the "draft label" mark and the footer counts them.

Work categories (ALWAYS)

Feature · Bug fix · Incident · Maintenance · Investigation · Experiment

Failure / waste types (ALWAYS)

Looping · Hedging · Wrong turn · Rework · Regression · Premature completion · Unauthorized action · Suboptimal path · Duplicate work · Blocked work · Missed requirement · Unnecessary escalation · Context re-read · Model thrash · Fan-out waste · Recovery after miss

Rework lock (ALWAYS): Rework lives only under failure_type — never as a work category. Keep category as the intended job type; set failure_type: Rework when rework occurred.

A line item can have a work category and a failure_type (e.g. Maintenance + Looping).

Cost model

Measured tokens come from Claude Code transcripts (~/.claude/projects/**/*.jsonl, assistant replies deduped on (message.id, requestId)); Codex transcripts are read the same way. Prices are OpenRouter public list prices per million tokens, fetched at run time and saved with the report.

agent_cost_i (per reply, micro-dollars) =
      input × price.input + output × price.output
    + cache_write_5m × price.cache_write + cache_write_1h × price.cache_write_1h
    + cache_read × price.cache_read                    # cache_write_1h = listed rate, else 2 × input

agent_estimated_usd        = Σ agent_cost_i over every reply in the window (matched or not)   # ESTIMATED, the headline
extrapolated_unmeasured    = observer tokens of sessions with no transcript here × (measured $ per observer token)
                           # EXTRAPOLATED (low confidence); "all sessions measured" when nothing remains
observer_note_taker_est    = note-taker (observer) tokens, deduped per reply, × its input list price
                           # priced separately, never agent cost, never in the headline
mistakes_estimated_usd     = Σ agent_cost_i over the same-session union of wasted turns, each turn once   # low figure
cost_per_completed_outcome = Σ attributed $ for status ∈ {shipped, completed} / count(those work items)
waste_rate                 = Σ wasted_cost / Σ attributed $     recovery_share = Σ recovery_cost / Σ attributed $

Line items are sessions: attributed_usd = estimated (transcript on this box) or extrapolated (no transcript). wasted_cost and recovery_cost come from the behavior pass (below), one union set for the ribbon, the line items and the mistakes line. The upper bound (redo windows plus project-wide fallback) stays in Details.

Unauthorized blocked: direct_cost $0, risk_exposure high, action_status blocked. Always keep risk_exposure non-dollar unless real cash/remediation is at stake — never invent risk dollars.

confidence — high when tokens + model + outcome are clear; medium when allocation across obs is judgmental; low when evidence is thin.

Unpriced models (not in the price list, or a negative "variable" price) are listed by name with their tokens and add nothing; they are never priced at zero silently.

Behavior metrics (heuristic until reviewed)

A second pass over the same transcripts tags every user turn human / bot / unknown (relayed agent prompts are never Alex's words), finds frustration episodes, and runs the pattern detectors from the Frustration Arc study: invented human gates, broke working things, wrong or expensive model, over-engineering, did something not asked, fake output, false "done", wrong tool or contact, bad outbound (incidents × recipients, never dollars), memory or rule loss, jargon, unclear cause, plus tool errors and hedging (Alex's definition: a caveat given when the answer was already available). Four summary tiles; everything else in Details. Every count is labeled heuristic until reviewed or classified. The optional classifier (--classify) is off by default, capped at $2.00 per run, uses only a regular inference OPENROUTER_API_KEY, and its spend is shown separately.

Money labeling (ALWAYS)

| Label | Meaning | |-------|---------| | Measured | Provider-reported spend from a sanctioned source (today: the OpenRouter per-key snapshot, shown with its UTC bucket label) | | Estimated | Measured transcript tokens × OpenRouter list price per MTok (the formula above) | | Extrapolated | Sessions with no transcript on this machine, from the observer-token ratio; always "low confidence" | | Unavailable | No measured value — write measured spend unavailable, never $0 spent for unknown | | risk_exposure | Severity / qualitative unless real cash is at stake (seats, refunds, SLA) |

cost_basis on each line item: estimated_usage, measured_provider, or extrapolated. Dollars are shown to two decimals everywhere ($109.25); every figure carries its label and basis.

Line-item schema

Each row in line-items.csv / report.json.line_items:

| Field | Notes | |-------|-------| | work_item_id | Stable id (WI-1 …) | | title | Human work-item title (ALWAYS invent at review; drafts use the session's first prompt) | | scope / project / worktree | Period, session or project; primary project; worktree when relevant | | session_ids / content_session_id | Mem session id(s) and the transcript join key | | status | shipped / completed / in_progress / abandoned / blocked | | category / work_category | One of the six work categories | | failure_type / failure_signals | Empty or failure types (Rework lives here) | | cost_measured | Number or unavailable | | cost_estimated / cost_extrapolated / attributed_usd | List-price estimate; extrapolation for sessions with no transcript; the one used | | wasted_cost / recovery_cost / productive_cost | USD from the behavior union (0 if none) | | risk_exposure | none / low / medium / high (non-dollar unless real cash) | | evidence_ids / summary_ids | Observation and summary IDs | | agent_tokens | {input, output, cache_write_5m, cache_write_1h, cache_read} | | observer_tokens | Note-taker tokens, Details only | | model / model_prices_usd_per_mtok | Dominant model and the prices used | | cost_basis | estimated_usage / measured_provider / extrapolated | | label_source | keyword (draft) / llm / human, with reviewed_by and reviewed_at_pt | | device | local, remote-N, or an export label such as mac | | behavior_counts | Pattern hits in this session | | recommended_action / confidence / date_pt / notes | One lever; high/medium/low; PT day; one plain sentence |

Deliverables (ALWAYS)

acr.py writes one directory containing:

  1. report.html — self-contained (inline CSS, no script, no external

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars97.1k
CategoryData
Updated1d ago
Forks8.6k

Languages

TypeScript

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions