anth-reliability-patterns
'Implement reliability patterns for Claude API: circuit breakers,
Install / Use
npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-reliability-patternsInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Our assessment of anth-reliability-patterns
anth-reliability-patterns scores 90/100 on our quality scale, 1035th of 3,845 Development & Engineering skills we index (top 27%).
Its SKILL.md is 7.7 KB long, well organised into 17 sections with 4 code examples: a thorough specification that gives an agent plenty to work with.
With 2,785 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 6 days ago, so anth-reliability-patterns is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
anth-reliability-patterns compared with similar skills
All 4 of these similar skills score higher than anth-reliability-patterns; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| anth-reliability-patterns (this skill)by jeremylongshore | 90 | 2.8k | 6d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 86.3k | 14d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.1k | today | CLAUDE.md |
| ai-job-searchby MadsLorentzen | 100 | 44.5k | today | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.7k | 4d ago | CLAUDE.md |
Frequently asked questions
- How do I install anth-reliability-patterns?
- Run
npx skills add jeremylongshore/tons-of-skills-marketplace --skill anth-reliability-patterns. The install tabs above show the steps for each supported agent. - Which AI agents does anth-reliability-patterns work with?
- It is written for Claude Code, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is anth-reliability-patterns safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is anth-reliability-patterns still maintained?
- The repository was last updated 6 days ago, so anth-reliability-patterns is actively maintained.
Skill content
View source on GitHubname: anth-reliability-patterns description: 'Implement reliability patterns for Claude API: circuit breakers,
graceful degradation, idempotency, and fallback strategies.
Trigger with phrases like "anthropic reliability", "claude circuit breaker",
"claude fallback", "anthropic fault tolerance".
' allowed-tools: Read, Write, Edit, Grep version: 1.7.0 license: MIT author: Jeremy Longshore jeremy@intentsolutions.io tags:
- saas
- ai
- anthropic compatibility: Designed for Claude Code
Anthropic Reliability Patterns
Overview
Production reliability patterns for Claude API: circuit breaker (prevent cascading failures), graceful degradation (serve fallbacks), idempotency (safe retries), and timeout management.
Circuit Breaker
import time
from enum import Enum
class CircuitState(Enum):
CLOSED = "closed" # Normal operation
OPEN = "open" # Failing, reject requests
HALF_OPEN = "half_open" # Testing recovery
class ClaudeCircuitBreaker:
def __init__(self, failure_threshold: int = 5, recovery_timeout: int = 60):
self.state = CircuitState.CLOSED
self.failures = 0
self.threshold = failure_threshold
self.recovery_timeout = recovery_timeout
self.last_failure_time = 0.0
def call(self, func, *args, **kwargs):
if self.state == CircuitState.OPEN:
if time.time() - self.last_failure_time > self.recovery_timeout:
self.state = CircuitState.HALF_OPEN
else:
raise Exception("Circuit breaker OPEN — Claude API unavailable")
try:
result = func(*args, **kwargs)
if self.state == CircuitState.HALF_OPEN:
self.state = CircuitState.CLOSED
self.failures = 0
return result
except Exception as e:
self.failures += 1
self.last_failure_time = time.time()
if self.failures >= self.threshold:
self.state = CircuitState.OPEN
raise
# Usage
breaker = ClaudeCircuitBreaker(failure_threshold=5, recovery_timeout=60)
def safe_claude_call(prompt: str) -> str:
try:
return breaker.call(
client.messages.create,
model="claude-sonnet-4-20250514",
max_tokens=1024,
messages=[{"role": "user", "content": prompt}]
).content[0].text
except Exception:
return "AI assistant is temporarily unavailable."
Graceful Degradation
import anthropic
def complete_with_fallback(prompt: str) -> str:
"""Try Sonnet → Haiku → cached response → static fallback."""
models = ["claude-sonnet-4-20250514", "claude-haiku-4-20250514"]
for model in models:
try:
msg = client.messages.create(
model=model,
max_tokens=1024,
messages=[{"role": "user", "content": prompt}]
)
return msg.content[0].text
except anthropic.RateLimitError:
continue # Try cheaper model
except anthropic.APIStatusError:
continue # Try next model
# All models failed — return cached or static response
cached = cache.get(f"claude:{hash(prompt)}")
if cached:
return f"[Cached response] {cached}"
return "Our AI assistant is temporarily unavailable. Please try again in a few minutes."
Idempotent Requests
import hashlib
import json
class IdempotentClaude:
def __init__(self):
self.client = anthropic.Anthropic()
self.cache = {} # Use Redis in production
def create_message(self, idempotency_key: str | None = None, **kwargs) -> str:
# Generate deterministic key from request params if not provided
if not idempotency_key:
idempotency_key = hashlib.sha256(
json.dumps(kwargs, sort_keys=True, default=str).encode()
).hexdigest()
# Return cached result for duplicate requests
if idempotency_key in self.cache:
return self.cache[idempotency_key]
msg = self.client.messages.create(**kwargs)
result = msg.content[0].text
self.cache[idempotency_key] = result
return result
Timeout Configuration
# Layer timeouts for defense-in-depth
client = anthropic.Anthropic(
timeout=60.0, # SDK-level timeout (covers connect + read)
max_retries=3, # Auto-retry on 429/5xx
)
# Per-request timeout override
msg = client.messages.create(
model="claude-haiku-4-20250514",
max_tokens=64,
messages=[{"role": "user", "content": "Quick question"}],
timeout=10.0 # Override for fast operations
)
Reliability Checklist
- [ ] Circuit breaker prevents cascading failures
- [ ] Graceful degradation serves fallback responses
- [ ] Idempotency keys prevent duplicate processing
- [ ] Timeouts configured at SDK and application level
- [ ] Health check probes API connectivity
- [ ] Retry logic uses exponential backoff (SDK default)
- [ ] Rate limit headers monitored for pre-emptive throttling
Prerequisites
- Define an approved model/workspace policy, request timeout, retry and circuit thresholds, fallback behavior, idempotency store, and rollback owner.
- Exercise the controls in a sandbox using synthetic prompts and a no-op downstream sink before enabling production traffic.
- Permit telemetry to contain only correlation IDs, status classes, model IDs, token counts, latency, breaker state, and aggregate fallback counts. Never store prompts, completions, credentials, or sensitive tool arguments.
Instructions
- Validate request scope and estimate budget before the call; reject unapproved models, destinations, or data classes before invoking the API.
- Apply bounded retries only to transient, repeat-safe failures. Coordinate backoff and breaker state across instances so a provider incident does not create a retry storm.
- Generate an application idempotency key from a stable request identity, not from an unredacted prompt. Cache only completed, policy-approved results and never deduplicate operations with unreviewed side effects.
- Route to an explicitly approved fallback or a static response when the breaker opens. A fallback must preserve authorization and data-handling rules rather than silently widening scope.
- Run a canary after configuration changes, compare error/latency/fallback metrics, and restore the prior configuration if thresholds or data controls regress. Expire test fixtures and temporary cache entries.
Output
Return a reliability receipt with correlation ID, breaker transition, retry attempts and reasons, fallback route, idempotency outcome, timeout, aggregate token/cost counters, canary result, rollback reference, and cleanup status. Keep content and secret fields redacted.
Error Handling
- Do not retry validation, authentication, permission, or policy failures; surface a stable operator-safe error and preserve the original request ID.
- When all fallbacks fail, fail closed with a bounded user-facing message and queue only work that has an explicit retention and replay policy.
- If a cached response is stale, scope-mismatched, or missing its policy version, discard it and use the static fallback.
- If duplicate suppression or breaker state is unavailable, stop new nonessential traffic rather than issuing uncoordinated retries.
Examples
In a sandbox, inject five synthetic 529 responses, confirm the breaker opens, then allow one half-open probe. Record retries=bounded; fallback=static; duplicate_writes=0; canary=pass; rollback=not-needed, with no prompt or completion in the receipt.
Resources
Next Steps
For policy guardrails, see anth-policy-guardrails.
Related Skills
Agent-Reach
86.3kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.1kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
ai-job-search
44.5kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.7kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
