auto-review-loop
Autonomous multi-round research review loop. In Copilot CLI it defaults to the native complementary rubber-duck subagent with host-event model evidence; elsewhere it uses Codex, while explicit external reviewer overrides remain available.
Install / Use
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill auto-review-loopInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Education & ResearchSupported Platforms
Our assessment of auto-review-loop
auto-review-loop scores 89/100 on our quality scale, 77th of 255 Education & Research skills we index (top 31%).
Its SKILL.md is 68 KB long, well organised into 52 sections with 17 code examples: long enough that it reads more like full documentation than a focused instruction file, which agents can find harder to follow.
With 16,644 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 9 days ago, so auto-review-loop is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
auto-review-loop compared with similar skills
All 4 of these similar skills score higher than auto-review-loop; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| auto-review-loop (this skill)by wanshuiyin | 89 | 16.6k | 9d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 85.8k | 12d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.0k | 1d ago | CLAUDE.md |
| rufloby ruvnet | 100 | 73.4k | today | CLAUDE.md |
| last30days-skillby mvanhorn | 100 | 63.0k | today | CLAUDE.md |
Frequently asked questions
- How do I install auto-review-loop?
- Run
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill auto-review-loop. The install tabs above show the steps for each supported agent. - Which AI agents does auto-review-loop work with?
- It is written for GitHub Copilot and OpenAI Codex, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is auto-review-loop safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is auto-review-loop still maintained?
- The repository was last updated 9 days ago, so auto-review-loop is actively maintained.
Skill content
View source on GitHubname: auto-review-loop description: Autonomous multi-round research review loop. In Copilot CLI it defaults to the native complementary rubber-duck subagent with host-event model evidence; elsewhere it uses Codex, while explicit external reviewer overrides remain available. Implements fixes and re-reviews until a policy-approved positive assessment or max rounds is reached. argument-hint: "[topic-or-scope]" allowed-tools: Bash(*), Read, Grep, Glob, Write, Edit, Skill, Task, mcp__codex__codex, mcp__codex__codex-reply, mcp__manual_review__review, mcp__manual_review__review_reply
Auto Review Loop: Autonomous Research Improvement
🔒 Do not wrap this skill in
/loop,/schedule, orCronCreate. It already loops internally (review → fix → re-review) and the reviewer carries round-to-round memory in onethreadId(codex-reply). An external timer re-enters from the top each tick — freshthreadId, reviewer memory reset — firing the verdict on wall-clock time instead of on artifact change: zero new signal, full token cost. If you want to schedule something, schedule the external wait that precedes it (experiments done → then run this once). Seeshared-references/external-cadence.md.
Autonomously iterate: review → implement fixes → re-review, until an independent reviewer gives a policy-approved positive assessment or MAX_ROUNDS is reached.
Context: $ARGUMENTS
Constants
- MAX_ROUNDS = 4
- POSITIVE_THRESHOLD: score >= 6/10 AND verdict ∈ {"ready", "almost"} — both must hold. This matches the operative Phase-E STOP CONDITION exactly; the verdict vocabulary is {"ready", "almost", "not ready"} (a high score with a "not ready" verdict does NOT stop the loop). Earlier wording here used
orand a stale verdict set ("accept"/"sufficient"/"ready for submission") — that was an internal inconsistency; theANDform is authoritative. - REVIEW_DOC:
review-stage/AUTO_REVIEW.md(cumulative log) (fall back to./AUTO_REVIEW.mdfor legacy projects) - REVIEWER_MODEL =
gpt-6-astra— Default model for the Codex backend. Must be an OpenAI model (e.g.,gpt-6-astra,o3,gpt-4o). Manual backend uses a model the user chooses — it must be a recognized model from a different family (OpenAI, Anthropic, Google, DeepSeek, Moonshot/Kimi, Qwen). - REVIEWER_BACKEND — With no reviewer directive, start as
auto; Step -1 runs exactly one two-call native marker/challenge probe for the first review. A bound Copilot CLI root session usescopilot-native(built-in complementaryrubber-ducksubagent); an unbound/non-Copilot host keeps the existingcodexdefault. Explicit— reviewer: codex,oracle-pro,agy, ormanualbypasses the probe and selects that external backend. Explicit— reviewer: copilotretains the compatibilitycopilot --agentdrive mode and its later Codex/manual finalizer. The native path gets both actual model IDs from host session events; it never needsCOPILOT_CLIor caller-provided--executor-model. Seeshared-references/reviewer-routing.md. - OUTPUT_DIR =
review-stage/— All review-stage outputs go here. Create the directory if it doesn't exist. - HUMAN_CHECKPOINT = false — When
true, pause after each round's review (Phase B) and present the score + weaknesses to the user. Wait for user input before proceeding to Phase C. The user can: approve the suggested fixes, provide custom modification instructions, skip specific fixes, or stop the loop early. Whenfalse(default), the loop runs fully autonomously. - COMPACT = false — When
true, (1) readEXPERIMENT_LOG.mdandfindings.mdinstead of parsing full logs on session recovery, (2) append key findings tofindings.mdafter each round. - REVIEWER_DIFFICULTY = medium — Controls how adversarial the reviewer is. Three levels:
medium(default): Current behavior — MCP-based review, the executor controls what context the reviewer sees.hard: Adds Reviewer Memory (the reviewer tracks its own suspicions across rounds) + Debate Protocol (the executor can rebut, the reviewer rules).nightmare: Everything inhard+ Codex exec reviewer reads the repo directly viacodex exec(the executor cannot filter what the reviewer sees) + Adversarial Verification (the reviewer independently checks if code matches claims).
- RENDER_HTML = true — When
true(default), auto-renderreview-stage/AUTO_REVIEW.mdto HTML on loop termination via/render-html. Uses--no-review(the loop itself IS the cross-model review; the HTML is a structural conversion). Setfalseto skip, or pass— render html: false.
⚠️ Nightmare + Manual incompatibility: If
REVIEWER_BACKEND = manualandREVIEWER_DIFFICULTY = nightmare, STOP with: "difficulty: nightmare requires Codex CLI / codex exec and is not compatible with --reviewer: manual. Use difficulty: hard, or switch reviewer to codex."
💡 Override:
/auto-review-loop "topic" — compact: true, human checkpoint: true, difficulty: hard
Reviewer Calling Convention
When calling the reviewer, branch on REVIEWER_BACKEND:
If no --reviewer: directive was supplied:
Set REVIEWER_BACKEND to auto. At Step -1 of the first round, resolve
copilot_native_evidence.py using the canonical four-layer helper chain.
Generate a fresh binding <run_id>_r<round>_review_<8-random-hex> and invoke
marker, wait, then invoke challenge as two distinct root Bash calls.
Put the literal binding and concrete resolved helper path in both calls;
Copilot Bash calls do not share variables. If the challenge binds, set
REVIEWER_BACKEND to copilot-native and use that same challenge for the
first review. Do not issue a second activation challenge in Phase A. If it
exits 3 because no current Copilot root session is bound, use codex.
Explicit reviewer directives bypass this probe. If the helper is missing,
native acceptance is unavailable; use Codex only if that external backend
is positively available, otherwise emit REVIEW_UNAVAILABLE.
If REVIEWER_BACKEND = copilot-native:
Read the challenge nonce and host-reported executor model. Invoke the host's
native task tool with agent_type: rubber-duck; do not start a subprocess
and do not specify a reviewer model. The prompt contains the exact standalone
ARIS_REVIEW_NONCE=<nonce> line, artifact/diff paths, the output contract,
and (round 2+) review-stage/REVIEWER_MEMORY.md. It contains no executor
summary or fix narrative. After the task completes, invoke
copilot_native_evidence.py verify to create the evidence and raw-response
artifacts. The verifier must observe one successful linked rubber-duck
lifecycle and known, different host-reported model families.
Pass the evidence to both review_gate.py --native-evidence and
save_trace.sh --backend copilot-native --native-evidence. A qualifying
native positive may stop directly; no external finalizer is needed. A native
negative continues with a fresh marker/challenge/subagent next round. Every
verdict-bearing native call—including a hard-mode rebuttal ruling—gets one
unique <run_id, round, purpose> artifact set and exactly one challenge.
Missing, same/unknown-family, malformed, stale, or mismatched evidence is
never a verdict. If native complementary dispatch is unavailable, fall back
only to a positively available opposite-family backend: Anthropic/Google
executor → Codex; OpenAI executor → manual with a reported non-OpenAI model.
Otherwise emit REVIEW_UNAVAILABLE. Full protocol:
shared-references/reviewer-routing.md.
If REVIEWER_BACKEND = copilot:
Require --executor-model: if not provided → emit REVIEW_UNAVAILABLE.
Determine executor family from --executor-model (see reviewer-routing.md).
Router picks opposite-family profile:
- executor_family=openai → profile="aris-reviewer-claude" (anthropic)
- executor_family=anthropic → profile="aris-reviewer-openai" (openai)
- executor_family=google → profile="aris-reviewer-openai" (openai, default cross)
- executor_family=unknown →
REVIEW_UNAVAILABLE(fail closed). Verify the profile file exists at.github/agents/<profile>.agent.md. If missing →REVIEW_UNAVAILABLE. Read itsmodel:field intoREVIEWER_MODEL, derivereviewer_familyfrom that model string, and verify it differs fromexecutor_family. Pass the same value through subprocess--model; never trust a caller-supplied family label or profile-only pinning under an Auto session. Identity assurance:--executor-modelis caller-declared routing input, not runtime attestation. Recordexecutor_model_source: caller-declared, the derivedfamily_relation, andindependence_verified: unverified. A pair of different model strings must never be promoted to independently verified. Capability gate:copilot --helpmust advertise--model,--effort, and--allow-tool; otherwise emitREVIEW_UNAVAILABLE. Use thecopilot --agentsubprocess (documented Copilot CLI form) with the selected profile,--model "$REVIEWER_MODEL",--effort xhigh, and--allow-tool=readfor each review call. Multi-round: each round is a freshcopilot --agentcall with the same profile; reviewer memory is carried viareview-stage/REVIEWER_MEMORY.mdartifact. IfcopilotCLI is unavailable →REVIEW_UNAVAILABLEfor that drive round; do not silently substitute another transport. A later positive Copilot verdict still requires the separately documented Codex/manual finalizer. Seeshared-references/reviewer-routing.mdfor the full copilot contract.
If REVIEWER_BACKEND = codex:
Use mcp__codex__codex for new review threads.
Use mcp__codex__codex-reply for follow-up rounds (reuse threadId).
If REVIEWER_BACKEND = manual:
Use mcp__manual_review__review for new review threads with:
prompt: [exact same prompt that would go to Codex]
config: {"model_reasoning_effort": "xhigh", "executor_model": "<actual executor model>", "require_reviewer_model": true}
Save the returned threadId.
Use mcp__manual_review__review_reply for follow-up rounds with:
threadId: [saved manual-review threadId]
prompt: [follow-up prompt]
config: {"model_reasoning_effort": "xhigh", "executor_model": "<actual executor model>", "require_reviewer_model": true}
A verdict-bearing manual response MUST begin with
Reviewer-Model: <exact-model-id>. Derive reviewer_family from that model
identity. Missing, unknown, or same-family identity cannot acquit; for a
mandatory escalation, emit REVIEW_UNAVAILABLE rather than guessing.
Prompt fidelity: the manual review task must be exactly the same text that Codex would receive; the transport may add only the required Reviewer-Model: response-format instruction.
Review tracing applies to every backend. Native traces are populated from the
revalidated host-event artifact rather than caller model declarations.
State Persistence (Compact Recovery)
Long-running loops may hit the context window limit, triggering automatic compaction. To survive this, persist state to review-stage/REVIEW_STATE.json after each round:
{
"run_id": "run_20260713_a1b2c3d4",
"round": 2,
"threadId": null,
"reviewer_profile": "rubber-duck",
"reviewer_backend": "copilot-native",
"executor_model": "claude-sonnet-4.6",
"executor_model_source": "host-session-event",
"executor_family": "anthropic",
"requested_reviewer_model": null,
"reported_reviewer_model": "gpt-5.5",
"reviewer_model_source": "host-session-event",
"reviewer_family": "openai",
"family_relation": "different",
"identity_assurance": "host_event_verified",
"independence_verified": true,
"native_evidence_id": "cne_0123456789abcdef0123456789abcdef",
"native_evidence_path": "review-stage/COPILOT_NATIVE_run_20260713_
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
85.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.0kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
ruflo
73.4k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
last30days-skill
63.0kAI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
