research-refine
Turn a vague research direction into a problem-anchored, elegant, frontier-aware, implementation-oriented method plan via iterative GPT-6-Astra review
Install / Use
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill research-refineInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AI & Machine LearningSupported Platforms
Our assessment of research-refine
research-refine scores 98/100 on our quality scale, 34th of 794 AI & Machine Learning skills we index (top 5%).
Its SKILL.md is 33 KB long, well organised into 75 sections with 15 code examples: a thorough specification that gives an agent plenty to work with.
With 16,644 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 9 days ago, so research-refine is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
research-refine compared with similar skills
All 4 of these similar skills score higher than research-refine; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| research-refine (this skill)by wanshuiyin | 98 | 16.6k | 9d ago | SKILL.md |
| claude-memby thedotmack | 100 | 94.8k | today | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 85.8k | 12d ago | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 84.4k | 16d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.0k | 1d ago | CLAUDE.md |
Frequently asked questions
- How do I install research-refine?
- Run
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill research-refine. The install tabs above show the steps for each supported agent. - Which AI agents does research-refine work with?
- It is written for OpenAI Codex, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is research-refine safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is research-refine still maintained?
- The repository was last updated 9 days ago, so research-refine is actively maintained.
Skill content
View source on GitHubname: research-refine description: 'Turn a vague research direction into a problem-anchored, elegant, frontier-aware, implementation-oriented method plan via iterative GPT-6-Astra review. Use when the user says "refine my approach", "帮我细化方案", "decompose this problem", "打磨idea", "refine research plan", "细化研究方案", or wants a concrete research method that stays simple, focused, and top-venue ready instead of a vague or overbuilt idea.' allowed-tools: Bash(*), Read, Write, Edit, Grep, Glob, WebSearch, WebFetch, mcp__codex__codex, mcp__codex__codex-reply
Research Refine: Problem-Anchored, Elegant, Frontier-Aware Plan Refinement
Refine and concretize: $ARGUMENTS
Overview
Use this skill when the research problem is already visible but the technical route is still fuzzy. The goal is not to produce a bloated proposal or a benchmark shopping list. The goal is to turn a vague direction into a problem -> focused method -> minimal validation document that is concrete enough to implement, elegant enough to feel paper-worthy, and current enough to resonate in the foundation-model era.
Four principles dominate this skill:
- Do not lose the original problem. Freeze an immutable Problem Anchor and reuse it in every round.
- The smallest adequate mechanism wins. Prefer the minimal intervention that directly fixes the bottleneck.
- One paper, one dominant contribution. Prefer one sharp thesis plus at most one supporting contribution.
- Modern leverage is a prior, not a decoration. When LLM / VLM / Diffusion / RL / distillation / inference-time scaling naturally fit the bottleneck, use them concretely. Do not bolt them on as buzzwords.
User input (PROBLEM + vague APPROACH)
-> Phase 0 (Claude): Freeze Problem Anchor
-> Phase 1 (Claude): Scan grounding papers -> identify technical gap -> choose the sharpest route -> write focused proposal
-> Phase 2 (Codex/GPT-6-Astra): Review for fidelity, specificity, contribution quality, and frontier leverage
-> Phase 3 (Claude): Anchor check + simplicity check -> revise method -> rewrite full proposal
-> Phase 4 (Codex, same thread): Re-evaluate revised proposal
-> Repeat Phase 3-4 until OVERALL SCORE >= 9 or MAX_ROUNDS reached
-> Phase 5: Save full history to refine-logs/
-> Optional handoff: /experiment-plan for a detailed execution-ready experiment roadmap
Constants
- REVIEWER_MODEL =
gpt-6-astra— Reviewer model used via Codex MCP. - MAX_ROUNDS = 5 — Maximum review-revise rounds.
- SCORE_THRESHOLD = 9 — Minimum overall score to stop.
- OUTPUT_DIR =
refine-logs/— Directory for round files and final report. - MAX_LOCAL_PAPERS = 15 — Maximum local papers/notes to scan for grounding.
- MAX_CORE_EXPERIMENTS = 3 — Default cap for core validation blocks inside this skill.
- MAX_PRIMARY_CLAIMS = 2 — Soft cap for paper-level claims. Prefer one dominant claim plus one supporting claim.
- MAX_NEW_TRAINABLE_COMPONENTS = 2 — Soft cap for genuinely new trainable pieces. Exceed only if the paper breaks otherwise.
Override via argument if needed, e.g.
/research-refine "problem | approach" -- max rounds: 3, threshold: 9.
State Persistence (Checkpoint Recovery)
Long-running refinement sessions may fail mid-way (e.g., API timeout, context compaction, or session interruption). To avoid losing completed work, persist state to refine-logs/REFINE_STATE.json after each phase boundary:
{
"phase": "review",
"round": 1,
"threadId": "019cd392-...",
"last_score": 6.5,
"last_verdict": "REVISE",
"status": "in_progress",
"timestamp": "2026-03-22T20:00:00"
}
Field definitions:
| Field | Values | Meaning |
|-------|--------|---------|
| phase | "anchor" / "proposal" / "review" / "refine" / "done" | Last completed phase |
| round | 0–MAX_ROUNDS | Current round number |
| threadId | string or null | Reviewer thread ID for codex-reply continuity |
| last_score | number or null | Most recent overall score from reviewer |
| last_verdict | string or null | Most recent verdict (READY / REVISE / RETHINK) |
| status | "in_progress" / "completed" | Loop status |
| timestamp | ISO 8601 | When state was last written |
Write rules:
- Write after each phase completes (not before). Overwrite each time — only the latest state matters.
- On completion (Phase 5 finished), set
"status": "completed".
Output Structure
refine-logs/
├── REFINE_STATE.json
├── round-0-initial-proposal.md
├── round-1-review.md
├── round-1-refinement.md
├── round-2-review.md
├── round-2-refinement.md
├── ...
├── REVIEW_SUMMARY.md
├── FINAL_PROPOSAL.md
├── REFINEMENT_REPORT.md
└── score-history.md
Every round-N-refinement.md must contain a full anchored proposal, not just incremental fixes.
Workflow
Initialization (Checkpoint Recovery)
Before starting any phase, check whether a previous run left a checkpoint:
-
Check for
refine-logs/REFINE_STATE.json:- If it does not exist → fresh start (proceed to Phase 0 normally)
- If it exists AND
statusis"completed"→ fresh start (delete state file, previous run finished) - If it exists AND
statusis"in_progress"ANDtimestampis older than 24 hours → fresh start (stale state from a killed/abandoned run — delete the file) - If it exists AND
statusis"in_progress"ANDtimestampis within 24 hours → resume
-
On resume, read the state file and recover context:
- Read all existing
refine-logs/round-*.mdfiles to restore prior work - Read
refine-logs/score-history.mdif it exists - Recover
threadIdfor reviewer thread continuity - Log to the user:
"Checkpoint found. Resuming after phase: {phase}, round: {round}." - Jump to the next phase based on the saved
phasevalue:
| Saved
phase| What was completed | Resume from | |---------------|-------------------|-------------| |"anchor"| Phase 0 done | Phase 1 (read anchor from round-0 context) | |"proposal"| Phase 1 done | Phase 2 (readround-0-initial-proposal.md) | |"review"| Phase 2 or 4 done | Phase 3 (read latestround-N-review.md) | |"refine"| Phase 3 done | Phase 4 (read latestround-N-refinement.md) | - Read all existing
-
On fresh start, ensure
refine-logs/directory exists and proceed to Phase 0.
Phase 0: Freeze the Problem Anchor
Before proposing anything, extract the user's immutable bottom-line problem. This anchor must be copied verbatim into every proposal and every refinement round.
Write:
- Bottom-line problem: What technical problem must be solved?
- Must-solve bottleneck: What specific weakness in current methods is unacceptable?
- Non-goals: What is explicitly not the goal of this project?
- Constraints: Compute, data, time, tooling, venue, deployment limits.
- Success condition: What evidence would make the user say "yes, this method addresses the actual problem"?
If later reviewer feedback would change the problem being solved, mark that as drift and push back or adapt carefully.
Checkpoint: Write refine-logs/REFINE_STATE.json with {"phase": "anchor", "round": 0, "threadId": null, "last_score": null, "last_verdict": null, "status": "in_progress", "timestamp": "<now>"}.
Phase 1: Build the Initial Proposal
Step 1.1: Scan Grounding Material
Check papers/ and literature/ first. Read only the relevant parts needed to answer:
- What mechanism do current methods use?
- Where exactly do they fail for this problem?
- Which recent LLM / VLM / Diffusion / RL era techniques are actually relevant here?
- What training objectives, representations, or interfaces are reusable?
- What details distinguish a real method from a renamed high-level idea?
If local material is insufficient, search recent top-venue/arXiv work online. Focus on method sections, training setup, and failure modes, not just abstracts.
Step 1.2: Identify the Technical Gap
Do not stop at generic research questions. Make the gap operational:
- Current pipeline failure point: where does the baseline break?
- Why naive fixes are insufficient: larger context, more data, prompting, memory bank, or stacking more modules.
- Smallest adequate intervention: what is the least additional mechanism that could plausibly fix the bottleneck?
- Frontier-native alternative: is there a more current route using foundation-model-era primitives that better matches the bottleneck?
- Core technical claim: what exact mechanism claim could survive top-venue scrutiny?
- Required evidence: what minimum proof is needed to defend that claim?
Step 1.3: Choose the Sharpest Route
Before locking the method, compare two candidate routes if both are plausible:
- Route A: Elegant minimal route — the smallest mechanism that directly targets the bottleneck.
- Route B: Frontier-native route — a more modern route that uses LLM / VLM / Diffusion / RL / distillation / inference-time scaling only if it gives a cleaner or stronger story.
Then decide:
- Which route is more likely to become a strong paper under the stated constraints?
- Which route has the cleaner novelty story relative to the closest work?
- Which route avoids contribution sprawl?
If both routes are weak, rethink the framing instead of combining them into a larger system by default.
Step 1.4: Concretize the Method First
The proposal must answer "how would we actually build this?" Prefer method detail over broad experimentation and prefer reuse over invention.
Cover:
- One-sentence method thesis: the single strongest mechanism claim.
- Contribution focus: one dominant contribution and at most one supporting contribution.
- Complexity budget: what is frozen or reused, what is new, and what tempting additions are intentionally excluded.
- System graph: modules, data flow, inputs, outputs.
- Representation design: what latent, embedding, plan token, reward signal, memory state, or alignment space is used?
- Training recipe: data source, supervision, pseudo-labeling, negatives, curriculum, losses, weighting, stagewise vs joint training.
- Inference path: how the trained components are used at test time and what signals flow where.
- Why the mechanism stays small: why a larger stack is unnecessary.
- Exact role of any frontier primitive: if you use an LLM / VLM / Diffusion / RL component, specify whether it acts as planner, teacher, critic, reward model, generator prior, search controller, or distillation source.
- Failure handling: what could go wrong and what fallback or diagnostic exists?
- Novelty and elegance argument: why this is more than naming a module and why the paper still looks focused.
If the method is still only described as "add a module" or "use a planner," it is not concrete enough.
Step 1.5: Design Minimal Claim-Driven Validation
Experiments exist to validate the method, not to dominate the document.
For each core claim, define the smallest strong experiment that can validate it:
- the claim being tested
- the necessary baseline or ablation
- the decisive metric
- the expected directional outcome
Additional rules:
- Ensure one experiment block directly supports the Problem Anchor.
- If complexity risk exists, include one simplification or deletion check.
- If a frontier primitive is central, include one necessity check showing why that choice matters.
- Default to 1-3 core experiment blocks and leave the full execution roadmap to
/experiment-plan.
Step 1.6: Write the Initial Proposal
Save to refine-logs/round-0-initial-proposal.md.
Use this structure:
# Research Proposal: [Title]
## Problem Anchor
- Bottom-line problem:
- Must-solve bottleneck:
- Non-goals:
- Constraints:
- Success condition:
## Technical Gap
[Why current metho
Truncated for display — read the full file on GitHub.
Related Skills
claude-mem
94.8kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
85.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Understand-Anything
84.4kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.0kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
