anth-observability
'Set up observability for Claude API integrations with metrics, logging,
Install / Use
npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-observabilityInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
OperationsSupported Platforms
Our assessment of anth-observability
anth-observability scores 88/100 on our quality scale, 305th of 548 Operations skills we index.
Its SKILL.md is 6.9 KB long, well organised into 16 sections with 3 code examples: a thorough specification that gives an agent plenty to work with.
With 2,785 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 6 days ago, so anth-observability is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
anth-observability compared with similar skills
All 4 of these similar skills score higher than anth-observability; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| anth-observability (this skill)by jeremylongshore | 88 | 2.8k | 6d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 86.3k | 14d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.1k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 84.6k | today | MCP Server |
| crawl4aiby unclecode | 100 | 84.5k | 5d ago | MCP Server |
Frequently asked questions
- How do I install anth-observability?
- Run
npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-observability. The install tabs above show the steps for each supported agent. - Which AI agents does anth-observability work with?
- It is written for Claude Code, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is anth-observability safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is anth-observability still maintained?
- The repository was last updated 6 days ago, so anth-observability is actively maintained.
Skill content
View source on GitHubname: anth-observability description: 'Set up observability for Claude API integrations with metrics, logging,
and alerting for latency, cost, errors, and token usage.
Trigger with phrases like "anthropic monitoring", "claude observability",
"anthropic metrics", "track claude usage", "claude dashboard".
' allowed-tools: Read, Write, Edit, Grep version: 1.7.0 license: MIT author: Jeremy Longshore jeremy@intentsolutions.io tags:
- saas
- ai
- anthropic compatibility: Designed for Claude Code
Anthropic Observability
Overview
Instrument Claude API calls with structured logging, Prometheus metrics, and cost tracking. Every API response includes usage data and rate limit headers — capture these for dashboards and alerting.
Structured Logging
import anthropic
import logging
import time
import json
logger = logging.getLogger("claude")
def create_with_logging(client: anthropic.Anthropic, **kwargs) -> anthropic.types.Message:
start = time.monotonic()
request_meta = {
"model": kwargs.get("model"),
"max_tokens": kwargs.get("max_tokens"),
"tool_count": len(kwargs.get("tools", [])),
"stream": kwargs.get("stream", False),
}
try:
response = client.messages.create(**kwargs)
duration_ms = int((time.monotonic() - start) * 1000)
logger.info(json.dumps({
"event": "claude.request",
"request_id": response._request_id,
"model": response.model,
"input_tokens": resp…[redacted],
"output_tokens": resp…[redacted],
"cache_read_tokens": getattr(response.usage, "cache_read_input_tokens", 0),
"stop_reason": response.stop_reason,
"duration_ms": duration_ms,
"content_blocks": len(response.content),
}))
return response
except anthropic.APIStatusError as e:
duration_ms = int((time.monotonic() - start) * 1000)
logger.error(json.dumps({
"event": "claude.error",
"status": e.status_code,
"error_type": getattr(e, "type", "unknown"),
"duration_ms": duration_ms,
"request_id": e.response.headers.get("request-id", "unknown"),
}))
raise
Prometheus Metrics
from prometheus_client import Counter, Histogram, Gauge
claude_requests = Counter(
"claude_requests_total", "Total Claude API requests",
["model", "stop_reason", "status"]
)
claude_latency = Histogram(
"claude_latency_seconds", "Claude API latency",
["model"], buckets=[0.5, 1, 2, 5, 10, 30, 60]
)
claude_tokens = Counter(
"claude_tokens_total", "Token usage",
["model", "direction"] # direction: input|output|cache_read
)
claude_cost = Counter(
"claude_cost_usd", "Estimated cost in USD",
["model"]
)
claude_rate_limit_remaining = Gauge(
"claude_rate_limit_remaining", "Remaining rate limit",
["dimension"] # dimension: requests|tokens
)
def track_metrics(response, duration: float):
model = response.model
claude_requests.labels(model=model, stop_reason=response.stop_reason, status="ok").inc()
claude_latency.labels(model=model).observe(duration)
claude_tokens.labels(model=model, direction="input").inc(response.usage.input_tokens)
claude_tokens.labels(model=model, direction="output").inc(response.usage.output_tokens)
# Cost estimation
pricing = {"claude-haiku-4-20250514": (0.80, 4.0), "claude-sonnet-4-20250514": (3.0, 15.0)}
rates = pricing.get(model, (3.0, 15.0))
cost = (response.usage.input_tokens * rates[0] + response.usage.output_tokens * rates[1]) / 1e6
claude_cost.labels(model=model).inc(cost)
Key Metrics Dashboard
| Metric | Description | Alert Threshold |
|--------|-------------|-----------------|
| claude_requests_total{status="error"} | Error count | > 5% of total |
| claude_latency_seconds p99 | Tail latency | > 10s |
| claude_cost_usd daily | Daily spend | > 80% budget |
| claude_rate_limit_remaining{dimension="requests"} | RPM headroom | < 10% remaining |
| claude_tokens_total{direction="output"} rate | Output throughput | Spike detection |
Usage API (Server-Side)
# Anthropic's Usage & Cost API for billing reconciliation
# GET https://api.anthropic.com/v1/usage
# Returns daily token usage and cost per model
Error Handling
| Observability Gap | Risk | Fix |
|-------------------|------|-----|
| No request_id logged | Can't debug with support | Capture response._request_id |
| Missing cost tracking | Budget surprise | Track per-request cost |
| No latency histogram | Can't spot slow queries | Add Prometheus/Datadog histograms |
Prerequisites
- Define SLOs, alert owners, budget and rate-limit thresholds, approved metric labels, and retention rules for telemetry.
- Configure authenticated server-side access through a secret manager and use a sandbox workspace with synthetic requests to verify instrumentation.
- Establish a redaction/filter policy before enabling logs, traces, dashboards, or usage reconciliation; prompts, responses, secrets, and personal data are never telemetry fields.
Instructions
- Instrument the request boundary with request ID, model, status, stop reason, token aggregates, cache counters, and duration while excluding content and high-cardinality identifiers.
- Emit success and failure metrics for authentication, 4xx/5xx, 429, timeout, latency, spend, and remaining rate-limit headroom. Validate labels against an allowlist and cap cardinality.
- Test dashboards and alerts with synthetic success, timeout, rate-limit, permission, and malformed-response fixtures. Verify the alert path without sending live customer data.
- Reconcile usage through the approved authenticated server-side API on a bounded schedule, compare aggregate totals, and alert on unexplained divergence or budget breach.
- Canary telemetry changes, then promote with owner approval. If redaction, cardinality, or retention checks fail, disable the new sink, restore the prior configuration, and preserve only a redacted receipt.
Output
Produce an observability receipt containing instrumentation version, metric/label allowlist, synthetic test results, alert thresholds, aggregate usage/cost/latency/error outcomes, retention policy, canary scope, approval, and rollback reference. Exclude prompts, responses, API keys, user IDs, and raw exception bodies.
Examples
Send synthetic fixture-request-001 through a staging client and assert request_id_present=1; content_fields=0; labels_allowlisted=1; inject a synthetic 429 and verify the alert fires. Record telemetry=pass; retention=24h; rollback=metrics-v1 without recording the fixture text.
Resources
Next Steps
For incident response, see anth-incident-runbook.
Related Skills
Agent-Reach
86.3kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.1kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Scrapling
84.6k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
crawl4ai
84.5kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
