SkillAgentSearch skills...

meta-optimize

Analyze ARIS usage logs and propose optimizations to SKILL.md files, reviewer prompts, and workflow defaults. Outer-loop harness optimization inspired by Meta-Harness (Lee et al., 2026)

Install / Use

npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill meta-optimize

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

98/100

Category

Automation

Supported Platforms

OpenAI Codex

Our assessment of meta-optimize

meta-optimize scores 98/100 on our quality scale, 76th of 1,943 Automation skills we index (top 4%).

Its SKILL.md is 23 KB long, well organised into 32 sections with 7 code examples: a thorough specification that gives an agent plenty to work with.

With 16,644 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
18/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 9 days ago, so meta-optimize is actively maintained.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

meta-optimize compared with similar skills

All 4 of these similar skills score higher than meta-optimize; compare them before choosing.

SkillScoreStarsUpdatedFormat
meta-optimize (this skill)by wanshuiyin9816.6k9d agoSKILL.md
Agent-Reachby Panniantong10085.8k12d agoCLAUDE.md
headroomby headroomlabs-ai10074.0k1d agoCLAUDE.md
rufloby ruvnet10073.4ktodayCLAUDE.md
CowAgentby zhayujie10047.1ktodayCLAUDE.md

Frequently asked questions

How do I install meta-optimize?
Run npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill meta-optimize. The install tabs above show the steps for each supported agent.
Which AI agents does meta-optimize work with?
It is written for OpenAI Codex, as a SKILL.md file. Other agents that read the same format can often use it too.
Is meta-optimize safe to use?
It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is meta-optimize still maintained?
The repository was last updated 9 days ago, so meta-optimize is actively maintained.

name: meta-optimize description: "Analyze ARIS usage logs and propose optimizations to SKILL.md files, reviewer prompts, and workflow defaults. Outer-loop harness optimization inspired by Meta-Harness (Lee et al., 2026). Use when user says "优化技能", "meta optimize", "improve skills", "分析使用记录", or wants to optimize ARIS's own harness components based on accumulated experience." argument-hint: "[target-skill-or-all]" allowed-tools: Bash(*), Read, Grep, Glob, mcp__codex__codex, mcp__codex__codex-reply

Meta-Optimize: Outer-Loop Harness Optimization for ARIS

Analyze accumulated usage logs and propose optimizations for: $ARGUMENTS

Privilege boundary — this skill is a READ-ONLY PRODUCER

meta-optimize proposes; it does not land. The mutation of the skill corpus is the exclusive job of a separate, human-invoked skill: /meta-apply. This split is structural, not advisory — it is why a missed instruction cannot let this loop apply its own patch (the self-acquittal failure mode):

  • No Write/Edit tool. This skill cannot edit a SKILL.md / shared-reference / any corpus file with the frictionless mutators. Its only outputs are the REPORT and staged patch files, written under .aris/meta/ (a scratch area, never the corpus).
  • No apply step. There is no in-skill "apply the patch" path (see Step 6). The producer ends by staging approved patches for /meta-apply; a human must then invoke /meta-apply to land them. That human action is the landing gate.
  • Bash writes to the corpus are filtered, not impossible — be honest about the layers. What IS fully closed: the accidental / in-flow self-acquittal — this skill has no Write/Edit and no apply step, so an honest run cannot slip into editing the corpus. Defense-in-depth: install the corpus_write_guard PreToolUse hook (like meta_logging.json), which DENIES the common Bash shell-writes (>, tee, sed -i, cp/mv, touch, open(...,'w')) to corpus paths. This is a blacklist, NOT a complete sandbox — a deliberately obscured Bash write (git apply, patch, $var/absolute paths, language file APIs) is not all caught. Full structural prevention requires either removing this skill's Bash or an FS sandbox — over-built for a not-yet-load-bearing producer, so deferred to when the gate carries real auto-modification volume (a brick-3 trigger). The intended backstop against a deliberate write is detection, not prevention — a corpus change with no valid/current provenance stamp (content-hash mismatch) would be catchable in a pre-push integrity check — but that verifier is NOT yet built (provenance.py has content_hash but no integrity-check subcommand, and no pre-push hook runs one). So today the deliberate-write case is neither prevented nor actively detected; track the integrity verifier as a follow-up before this producer goes load-bearing. Its legitimate Bash writes go only to .aris/meta/.

See shared-references/acceptance-gate.md: a loop can DRIVE (propose, review) same-model, but the ACQUITTAL that lands a change must be cross-model (Step 4 jury) and the landing must be a separate human-gated act (/meta-apply).

Context

ARIS is a research harness — a system of skills, bridges, workflows, and artifact contracts that wraps around LLMs to orchestrate research. This skill implements a prototype outer loop that observes how the harness is used and proposes improvements to the harness itself (not to the research artifacts it produces).

Inspired by Meta-Harness (Lee et al., 2026): the key insight is that harness design matters as much as model weights, and harness engineering can be partially automated by logging execution traces and using them to guide improvements.

What This Skill Optimizes (Harness Components)

| Component | Example | Optimizable? | |-----------|---------|:---:| | SKILL.md prompts | Reviewer instructions, quality gates, step descriptions | Yes | | Default parameters | difficulty: medium, MAX_ROUNDS: 4, threshold: 6/10 | Yes | | Convergence rules | When to stop the review loop, retry counts | Yes | | Workflow ordering | Skill chain sequence within a workflow | Yes | | Artifact schemas | What fields go in EXPERIMENT_LOG.md, idea-stage/IDEA_REPORT.md | Cautious | | MCP bridge config | Which reviewer model, routing rules | No (infra) |

Not optimized: The research artifacts themselves (papers, code, experiments). That's what the regular workflows do.

Prerequisites

  1. Logging must be active. Copy templates/claude-hooks/meta_logging.json into your project's .claude/settings.json (or merge the hooks section).
  2. Sufficient data. At least 5 complete workflow runs logged in .aris/meta/events.jsonl. The skill will check and warn if insufficient.

Workflow

Step 0: Check Data Availability

EVENTS_FILE=".aris/meta/events.jsonl"
if [ ! -f "$EVENTS_FILE" ]; then
    echo "ERROR: No event log found at $EVENTS_FILE"
    echo "Enable logging first: copy templates/claude-hooks/meta_logging.json into .claude/settings.json"
    exit 1
fi

EVENT_COUNT=$(wc -l < "$EVENTS_FILE")
SKILL_INVOCATIONS=$(grep -c '"skill_invoke"' "$EVENTS_FILE" || echo 0)
SESSIONS=$(grep -c '"session_start"' "$EVENTS_FILE" || echo 0)

echo "📊 Event log: $EVENT_COUNT events, $SKILL_INVOCATIONS skill invocations, $SESSIONS sessions"

if [ "$SKILL_INVOCATIONS" -lt 5 ]; then
    echo "⚠️  Insufficient data (<5 skill invocations). Continue using ARIS normally and re-run later."
    exit 0
fi

# Bottleneck succession: what did the LAST cycle say was the limiting stage?
BOTTLENECK_LOG=".aris/meta/bottleneck_log.jsonl"
if [ -f "$BOTTLENECK_LOG" ]; then
    echo "🧭 Prior cycle's bottleneck: $(tail -1 "$BOTTLENECK_LOG")"
fi

If a prior bottleneck entry exists, open the report (Step 5) by stating whether that named bottleneck was resolved (and by which landed patches) and what it has now moved to — bottleneck SUCCESSION, not just existence, is the signal this ledger exists to carry.

Step 1: Analyze Usage Patterns

Read .aris/meta/events.jsonl and compute:

Frequency analysis:

  • Which skills are invoked most often?
  • Which slash commands do users type most?
  • What parameter overrides are most common? (These suggest bad defaults.)

Failure analysis:

  • Which tools fail most often? In which skills?
  • What error patterns repeat? (OOM, import, compilation, timeout)
  • How many auto-debug retries per workflow run?

Convergence analysis (for auto-review-loop):

  • Average rounds to reach threshold
  • Score trajectory shape (fast improvement? plateau? oscillation?)
  • Which review round catches the most critical issues?
  • Do users override difficulty mid-run?

Human intervention analysis:

  • Where do users interrupt with manual prompts during workflows?
  • What manual corrections do users make most? (These indicate skill gaps.)

Model-delta analysis (harness diet):

  • Has the session model (session_start events' model field) or the pinned reviewer model changed since a skill's SKILL.md was last touched? (git log -1 --format=%cs -- skills/<skill>/SKILL.md vs the model-bump date.)
  • A model bump is a trigger to re-read, not evidence by itself. For each reasoning-scaffolding step or worked example in that SKILL.md, a deletion proposal must cite TARGET-SPECIFIC evidence that the new model no longer needs it: a capability-specific release note, or repeated observed behavior in the event log (e.g. zero failures/interventions in the guarded step since the bump). "The model got newer" alone never justifies a deletion.
  • Never deletion candidates, regardless of model: privilege boundaries, acceptance/review gates, corpus- and provenance-integrity rules, output contracts, and safety checks. The diet targets model-compensation scaffolding only — a capability the new model has natively is pure overhead (context weight, drift surface, reading cost). A harness that only ever grows is a harness nobody is re-reading.

Trigger-rate analysis (optional, measured — not from the event log):

  • The event log shows which skills were USED, not which were WANTED-but-omitted — the omission failure mode (Claude Code passing over the right skill when the installed list is long) is invisible to it. tools/meta_opt/trigger_eval.py measures it directly: claude -p probes with paraphrased-intent queries run from a neutral cwd (so the realistic long installed corpus is loaded), scored as trigger / confusion(→which skill) / miss.
  • Run it when a specific skill is suspected of under- or mis-triggering, or as a before/after check around a description edit: python3 tools/meta_opt/trigger_eval.py --eval-file tools/meta_opt/trigger_evals.sample.json --skills <name> --samples 2
  • The confusion matrix is the signal, not just the rate: a query that keeps landing on a sibling skill means the two descriptions overlap on that intent — the fix is disambiguation, not "make the description pushier".
  • Measure-only, evidence not verdict. A low trigger rate is an INPUT to a Step-2 proposal (which lands only via /meta-apply), never a self-applied description rewrite. Trigger rate is model-dependent, so compare like with like (record the probe model) and treat it as a proxy — it measures selection under a query set, not the full long-list omission problem.

Present findings as a structured summary table.

Step 1.5: Name the Current Bottleneck

Synthesize the Step-1 analyses into one sentence naming the single most-limiting pipeline stage right now — e.g. "planning", "verification quality", "experiment execution reliability", "writing polish" — with the supporting evidence. The bottleneck always moves: when coding stops being the constraint, planning becomes it; when planning is solved, verification; when verification is automated, taste. This step exists to make the CURRENT constraint visible, so Step 2's ranked table reads as sub-fixes for one named constraint instead of scattered tweaks.

Append the verdict to the append-only ledger .aris/meta/bottleneck_log.jsonl (same never-mutate discipline as .aris/runs/<run_id>.iterations.jsonl):

mkdir -p .aris/meta
# json.dumps, NOT hand-interpolated shell strings: bottleneck/evidence are
# natural language — a stray quote must not break the JSONL (or the shell).
python3 - <<'PY'
import json, datetime
entry = {
    "ts": datetime.datetime.now().astimezone().isoformat(timespec="seconds"),
    "cycle": 3,
    "bottleneck": "verification quality",
    "evidence": "review rounds plateau at 6/10 while tool failures are rare",
    "top_patch_ids": ["P1", "P2"],
}
with open(".aris/meta/bottleneck_log.jsonl", "a", encoding="utf-8") as fh:
    fh.write(json.dumps(entry, ensure_ascii=False) + "\n")
PY

Never edit or delete prior lines — succession history is the point.

Step 2: Identify Optimization Targets

Based on Step 1, rank optimization opportunities by expected impact:

## Optimization Opportunities (ranked)

| # | Target | Signal | Proposed Change | Expected Impact |
|---|--------|--------|-----------------|-----------------|
| 1 | auto-review-loop default threshold | Users override to 7/10 in 60% of runs | Change default from 6/10 to 7/10 | Fewer manual overrides |
| 2 | experiment-bridge retry count | 40% of runs hit max retries on OOM | Add OOM-specific recovery (reduce batch size) | Fewer failed experiments |
| 3 | paper-write de-AI patterns | Users manually fix "delve" in 80% of runs | Add "delve" to default watchword list | Fewer manual edits |
| 4 | experiment-bridge Phase-2 hand-holding steps | Model bump (session_start model changed); scaffold untouched since 2 generations ago; zero tool_failures in the steps it 

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars16.6k
CategoryAutomation
Updated9d ago
Forks1.4k

Languages

Python

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions