Agent Cost Report
Believable agent cost report for any period, default the last 7 full days PT, not counting today. Measured tokens from Claude Code transcripts priced at OpenRouter list prices (ESTIMATED), measured provider spend when a sanctioned source exists, note-taker cost separate, Timing-style HTML/PDF plus r…
Install / Use
npx skills add thedotmack/claude-mem --skill agent-cost-reportInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Data & AnalyticsSupported Platforms
Our assessment of Agent Cost Report
Agent Cost Report scores 97/100 on our quality scale, 25th of 597 Data & Analytics skills we index (top 5%).
Its SKILL.md is 20 KB long, well organised into 21 sections with 1 code example: a thorough specification that gives an agent plenty to work with.
With 97,136 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated yesterday, so Agent Cost Report is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Agent Cost Report compared with similar skills
All 4 of these similar skills score higher than Agent Cost Report; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| Agent Cost Report (this skill)by thedotmack | 97 | 97.1k | 1d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 93.0k | 22d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.6k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.3k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 86.1k | today | MCP Server |
Frequently asked questions
- How do I install Agent Cost Report?
- Run
npx skills add thedotmack/claude-mem --skill "Agent Cost Report". The install tabs above show the steps for each supported agent. - Which AI agents does Agent Cost Report work with?
- It is written for Claude Code, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is Agent Cost Report safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is Agent Cost Report still maintained?
- The repository was last updated yesterday, so Agent Cost Report is actively maintained.
Skill content
View source on GitHubname: Agent Cost Report description: >- Believable agent cost report for any period, default the last 7 full days PT, not counting today. Measured tokens from Claude Code transcripts priced at OpenRouter list prices (ESTIMATED), measured provider spend when a sanctioned source exists, note-taker cost separate, Timing-style HTML/PDF plus report.json, line-items.csv, evidence.json. allowed-tools:
- Bash
- Read
- Write
- AskUserQuestion
- mcp__plugin_claude-mem_mcp-search__search
- mcp__plugin_claude-mem_mcp-search__timeline
- mcp__plugin_claude-mem_mcp-search__get_observations
Agent Cost Report
Claude-Mem / Claude Code skill. Runtime is the scripts/ pipeline (transcripts → tokens → dollars → Timing-style report) plus a progressive Mem Search review pass that confirms the drafted labels. The Notion draft is SPEC history only — never the product, never the runtime, never the ship vehicle.
Resolve the absolute directory containing this SKILL.md; all helper paths are relative to that directory. ${CLAUDE_SKILL_DIR} is the shortcut: python3 "${CLAUDE_SKILL_DIR}/scripts/acr.py" …. Python 3.9+ standard library only, with IANA timezone data for America/Los_Angeles; no pip installs. tzdata is the only exception to the no-pip-installs rule, and only when acr.py exits saying that timezone data is unavailable (stock Windows Python ships none): show the user the lines it printed and ask (AskUserQuestion) before installing it. On a yes, run the line for your shell (the plain quoted line in Bash or cmd, the PowerShell: line in PowerShell), then rerun the failed command. Setting PYTHONTZPATH instead needs no install. The look lives in scripts/acr/render.py, never here.
Purpose
Turn Claude-Mem activity into a manager-readable cost and failure report.
Product idea: a reusable skill that searches Claude-Mem via Mem Search, reconstructs real units of work, assigns cost and failure categories, and renders a printable report.
Insight north star
The headline is dollars, to two decimals, labeled. The dollars come from Claude Code transcripts (exact per-reply token usage) priced at OpenRouter public list prices, so they are ESTIMATED. Measured provider spend appears only when a sanctioned source gives it. Directly under the dollars: what the mistakes cost, what shipped, and both on one time axis.
Questions the report must answer
- What work was completed?
- What did each outcome cost?
- What was wasted through looping, hedging, wrong turns, rework, or poor routing?
- Were any unauthorized actions attempted?
- What should the manager change next?
Primary unit = cost per completed outcome (not cost per observation).
When to use
- "Agent cost report" / "cost per outcome" / "failure economics" / "was this session worth it" / "what did the agents cost this week"
- After a real Mem session dig when leadership needs outcome economics
- Sample / ship packs that need self-contained HTML + JSON + CSV + evidence
Memory dig mechanics: the claude-mem mem-search skill (progressive recall).
Default scope (ALWAYS)
- Unless the user names a specific session / range / project, the window is the last 7 full days in PT, not counting today:
end= the PT midnight that started today (exclusive),start=end − 7 days. The default window never contains a partial day (G3, Alex 2026-09-25). - Explicit windows:
--start YYYY-MM-DD --end YYYY-MM-DD(PT calendar days, end exclusive). An explicit--endlater than today marks the last day "partial, generated HH:MM PT". - One session:
--session <content_session_id>. One project plus a period:--project <name> --start … --end …(worktrees of the project are included). - The report Scope strip shows the PT range, and Details list every session id in scope.
Progressive Mem Search (ALWAYS)
Follow the claude-mem mem-search three layers. Keep spend light.
- Search — get an index of IDs (titles, types, token hints).
- Timeline — only around anchors you care about.
- Observations —
get_observationsfor the filtered IDs you will cite as evidence.
Recipe:
- Resolve scope (default: the last 7 full PT days).
- Search → collect IDs.
- Timeline for thin context only.
- Observations for intended / actual / outcome / waste / rework / blocked / unauthorized / status.
- Group into named work items + failure events.
- Calculate line-item costs.
- Render HTML + optional PDF (+ json/csv/evidence).
- Keep evidence IDs in the appendix — do not dump entire timelines into the main report.
ADHD process bullets:
- Search first → pick IDs → timeline only if context is thin → fetch only needed obs.
- Work-item titles are invented for managers ("Restore search after Chroma crash-loop"); observation titles stay evidence-only.
- Evidence appendix lists obs IDs + short titles; main sections stay outcome-first.
In this skill the search pass is the review step (see Recipe step 4): the pipeline drafts categories and failure types from keywords (label_source: keyword); the orchestrator confirms or changes each line item's category and failure_type from its cited evidence IDs and applies the result with review --apply. Items left unreviewed keep the "draft label" mark and the footer counts them.
Work categories (ALWAYS)
Feature · Bug fix · Incident · Maintenance · Investigation · Experiment
Failure / waste types (ALWAYS)
Looping · Hedging · Wrong turn · Rework · Regression · Premature completion · Unauthorized action · Suboptimal path · Duplicate work · Blocked work · Missed requirement · Unnecessary escalation · Context re-read · Model thrash · Fan-out waste · Recovery after miss
Rework lock (ALWAYS): Rework lives only under failure_type — never as a work category. Keep category as the intended job type; set failure_type: Rework when rework occurred.
A line item can have a work category and a failure_type (e.g. Maintenance + Looping).
Cost model
Measured tokens come from Claude Code transcripts (~/.claude/projects/**/*.jsonl, assistant replies deduped on (message.id, requestId)); Codex transcripts are read the same way. Prices are OpenRouter public list prices per million tokens, fetched at run time and saved with the report.
agent_cost_i (per reply, micro-dollars) =
input × price.input + output × price.output
+ cache_write_5m × price.cache_write + cache_write_1h × price.cache_write_1h
+ cache_read × price.cache_read # cache_write_1h = listed rate, else 2 × input
agent_estimated_usd = Σ agent_cost_i over every reply in the window (matched or not) # ESTIMATED, the headline
extrapolated_unmeasured = observer tokens of sessions with no transcript here × (measured $ per observer token)
# EXTRAPOLATED (low confidence); "all sessions measured" when nothing remains
observer_note_taker_est = note-taker (observer) tokens, deduped per reply, × its input list price
# priced separately, never agent cost, never in the headline
mistakes_estimated_usd = Σ agent_cost_i over the same-session union of wasted turns, each turn once # low figure
cost_per_completed_outcome = Σ attributed $ for status ∈ {shipped, completed} / count(those work items)
waste_rate = Σ wasted_cost / Σ attributed $ recovery_share = Σ recovery_cost / Σ attributed $
Line items are sessions: attributed_usd = estimated (transcript on this box) or extrapolated (no transcript). wasted_cost and recovery_cost come from the behavior pass (below), one union set for the ribbon, the line items and the mistakes line. The upper bound (redo windows plus project-wide fallback) stays in Details.
Unauthorized blocked: direct_cost $0, risk_exposure high, action_status blocked. Always keep risk_exposure non-dollar unless real cash/remediation is at stake — never invent risk dollars.
confidence — high when tokens + model + outcome are clear; medium when allocation across obs is judgmental; low when evidence is thin.
Unpriced models (not in the price list, or a negative "variable" price) are listed by name with their tokens and add nothing; they are never priced at zero silently.
Behavior metrics (heuristic until reviewed)
A second pass over the same transcripts tags every user turn human / bot / unknown (relayed agent prompts are never Alex's words), finds frustration episodes, and runs the pattern detectors from the Frustration Arc study: invented human gates, broke working things, wrong or expensive model, over-engineering, did something not asked, fake output, false "done", wrong tool or contact, bad outbound (incidents × recipients, never dollars), memory or rule loss, jargon, unclear cause, plus tool errors and hedging (Alex's definition: a caveat given when the answer was already available). Four summary tiles; everything else in Details. Every count is labeled heuristic until reviewed or classified. The optional classifier (--classify) is off by default, capped at $2.00 per run, uses only a regular inference OPENROUTER_API_KEY, and its spend is shown separately.
Money labeling (ALWAYS)
| Label | Meaning |
|-------|---------|
| Measured | Provider-reported spend from a sanctioned source (today: the OpenRouter per-key snapshot, shown with its UTC bucket label) |
| Estimated | Measured transcript tokens × OpenRouter list price per MTok (the formula above) |
| Extrapolated | Sessions with no transcript on this machine, from the observer-token ratio; always "low confidence" |
| Unavailable | No measured value — write measured spend unavailable, never $0 spent for unknown |
| risk_exposure | Severity / qualitative unless real cash is at stake (seats, refunds, SLA) |
cost_basis on each line item: estimated_usage, measured_provider, or extrapolated. Dollars are shown to two decimals everywhere ($109.25); every figure carries its label and basis.
Line-item schema
Each row in line-items.csv / report.json.line_items:
| Field | Notes |
|-------|-------|
| work_item_id | Stable id (WI-1 …) |
| title | Human work-item title (ALWAYS invent at review; drafts use the session's first prompt) |
| scope / project / worktree | Period, session or project; primary project; worktree when relevant |
| session_ids / content_session_id | Mem session id(s) and the transcript join key |
| status | shipped / completed / in_progress / abandoned / blocked |
| category / work_category | One of the six work categories |
| failure_type / failure_signals | Empty or failure types (Rework lives here) |
| cost_measured | Number or unavailable |
| cost_estimated / cost_extrapolated / attributed_usd | List-price estimate; extrapolation for sessions with no transcript; the one used |
| wasted_cost / recovery_cost / productive_cost | USD from the behavior union (0 if none) |
| risk_exposure | none / low / medium / high (non-dollar unless real cash) |
| evidence_ids / summary_ids | Observation and summary IDs |
| agent_tokens | {input, output, cache_write_5m, cache_write_1h, cache_read} |
| observer_tokens | Note-taker tokens, Details only |
| model / model_prices_usd_per_mtok | Dominant model and the prices used |
| cost_basis | estimated_usage / measured_provider / extrapolated |
| label_source | keyword (draft) / llm / human, with reviewed_by and reviewed_at_pt |
| device | local, remote-N, or an export label such as mac |
| behavior_counts | Pattern hits in this session |
| recommended_action / confidence / date_pt / notes | One lever; high/medium/low; PT day; one plain sentence |
Deliverables (ALWAYS)
acr.py writes one directory containing:
report.html— self-contained (inline CSS, no script, no external
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
93.0kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.6kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
CowAgent
47.3kOpen-source personal AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Scrapling
86.1k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
