lab:autoresearch
Self-improving loop for plugin skills. Reads program.md, proposes one mutation per iteration, evaluates against deterministic scorer, keeps improvements via git, reverts failures. Targets weakest skill+dimension. Use with /loop for overnight runs.
Install / Use
npx skills add oliver-kriska/claude-elixir-phoenix --skill autoresearchInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
MarketingSupported Platforms
Our assessment of lab:autoresearch
lab:autoresearch scores 87/100 on our quality scale, 276th of 615 Marketing skills we index (top 45%).
Its SKILL.md is 4.5 KB long, well organised into 18 sections with 5 code examples: a solid amount of guidance for an agent.
It has 560 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 2 days ago, so lab:autoresearch is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
lab:autoresearch compared with similar skills
All 4 of these similar skills score higher than lab:autoresearch; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| lab:autoresearch (this skill)by oliver-kriska | 87 | 560 | 2d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 89.8k | 18d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 12d ago | SKILL.md |
Frequently asked questions
- How do I install lab:autoresearch?
- Run
npx skills add oliver-kriska/claude-elixir-phoenix --skill "lab:autoresearch". The install tabs above show the steps for each supported agent. - Which AI agents does lab:autoresearch work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is lab:autoresearch safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is lab:autoresearch still maintained?
- The repository was last updated 2 days ago, so lab:autoresearch is actively maintained.
Skill content
View source on GitHubname: lab:autoresearch description: > Self-improving loop for plugin skills. Reads program.md, proposes one mutation per iteration, evaluates against deterministic scorer, keeps improvements via git, reverts failures. Targets weakest skill+dimension. Use with /loop for overnight runs. effort: high argument-hint: "[--skill NAME] [--strategy targeted|sweep|random] [--dry-run] [--max-iterations N]" disable-model-invocation: true
Autoresearch — Plugin Skill Self-Improvement
Iteratively improve plugin skills via the autoresearch pattern: propose one mutation -> eval -> keep/revert -> repeat.
Usage
/lab:autoresearch # Targeted: attack weakest skill+dimension
/lab:autoresearch --skill review # Focus on one skill
/lab:autoresearch --strategy sweep # Process all skills alphabetically
/lab:autoresearch --dry-run # Show what would change, don't commit
For overnight runs:
/loop 5m /lab:autoresearch --strategy sweep --max-iterations 200
Iron Laws
- ONE mutation per iteration — if description needs "and", split into two
- NEVER mutate read-only files — check program.md before every write
- EVAL is deterministic — always use the wrapper script, never LLM-judge
- REVERT on regression OR checks failure — no exceptions
- LOG every iteration — use
keeporrevertcommand (never skip) - CHECK ideas.md before proposing — don't rediscover known optimizations
Wrapper Script Commands
All eval/git/journal operations go through ONE script. Do NOT run these manually.
# Find the weakest skill+dimension
python3 lab/autoresearch/scripts/run-iteration.py target --strategy targeted
# Score a skill (before mutation, to get baseline)
python3 lab/autoresearch/scripts/run-iteration.py score <skill-name>
# After mutation: score + checks + compare → verdict (KEEP or REVERT)
python3 lab/autoresearch/scripts/run-iteration.py eval <skill-name>
# Act on verdict:
python3 lab/autoresearch/scripts/run-iteration.py keep <skill> <dim> <old> <new> \
--desc "what changed" --asi '{"hypothesis": "why", "mechanism": "how"}'
python3 lab/autoresearch/scripts/run-iteration.py revert <skill> <dim> <old> <new> \
--desc "what was attempted" --asi '{"hypothesis": "why", "regression": "what broke", "avoid": "do not retry this"}'
# Check overall progress
python3 lab/autoresearch/scripts/run-iteration.py status
Core Loop (ONE iteration)
Step 1: Read State
- Read
lab/autoresearch/program.md(goals, mutable surface, rules) - Read
lab/autoresearch/ideas.mdif it exists (deferred optimizations) - Run:
python3 lab/autoresearch/scripts/run-iteration.py status
Step 2: Select Target
Run: python3 lab/autoresearch/scripts/run-iteration.py target --strategy targeted
Parse the JSON: skill, dimension, failing_checks. If all_perfect → STOP.
Step 3: Read + Propose
- Read target SKILL.md and its references/ listing
- Read eval definition from
lab/eval/evals/{skill}.json - Check
ideas.mdfor deferred ideas about this skill - Check recent journal entries for prior failures on this skill (avoid repeats)
- Consult
${CLAUDE_SKILL_DIR}/references/mutation-strategies.md - Propose exactly ONE change targeting the failing checks
Step 4: Apply + Evaluate
- Apply the mutation via Edit tool
- Run:
python3 lab/autoresearch/scripts/run-iteration.py eval <skill-name> - Parse JSON → check
verdictfield
Step 5: Keep or Revert
If verdict is KEEP:
python3 lab/autoresearch/scripts/run-iteration.py keep <skill> <dim> <old> <new> \
--desc "..." --asi '{"hypothesis": "...", "mechanism": "..."}'
If verdict is REVERT:
python3 lab/autoresearch/scripts/run-iteration.py revert <skill> <dim> <old> <new> \
--desc "..." --asi '{"hypothesis": "...", "regression": "...", "avoid": "..."}'
Step 6: Ideas Backlog
If during analysis you discovered a promising optimization you can't act on now:
- Append it to
lab/autoresearch/ideas.mdas a bullet - On next resume: prune stale/tried ideas, experiment with the rest
Step 7: Continue or Stop
- All targets >= 0.95? Print "AUTORESEARCH_COMPLETE"
- Max iterations reached? Print "AUTORESEARCH_COMPLETE"
- 50 consecutive discards? Print "AUTORESEARCH_STUCK"
- Otherwise: immediately start Step 1 again
References
${CLAUDE_SKILL_DIR}/references/mutation-strategies.md— mutation type catalog${CLAUDE_SKILL_DIR}/references/state-management.md— git protocol, journalinglab/autoresearch/program.md— research agenda (read every iteration)
Related Skills
Agent-Reach
89.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
