harness-eval
Evaluate a repo agent harness (AGENTS.md, rules, skills, skill refs) for broken paths/commands, redundant instructions, and usefulness using a stack-agnostic dual-judge protocol with planted traps.
Install / Use
npx skills add tech-leads-club/agent-skills --skill harness-evalInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Tags
Our assessment of harness-eval
harness-eval scores 96/100 on our quality scale, 187th of 3,044 Development & Engineering skills we index (top 7%).
Its SKILL.md is 15 KB long, well organised into 36 sections with 8 code examples: a thorough specification that gives an agent plenty to work with.
With 6,832 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 7 days ago, so harness-eval is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-09-28. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
harness-eval compared with similar skills
All 4 of these similar skills score higher than harness-eval; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| harness-eval (this skill)by tech-leads-club | 96 | 6.8k | 7d ago | SKILL.md |
| ai-job-searchby MadsLorentzen | 100 | 44.2k | today | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.7k | 1d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 5d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 5d ago | SKILL.md |
Frequently asked questions
- How do I install harness-eval?
- Run
npx skills add tech-leads-club/agent-skills --skill harness-eval. The install tabs above show the steps for each supported agent. - Which AI agents does harness-eval work with?
- It is written for OpenAI Codex, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is harness-eval safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is harness-eval still maintained?
- The repository was last updated 7 days ago, so harness-eval is actively maintained.
Skill content
View source on GitHubname: harness-eval description: "Evaluate a repo agent harness (AGENTS.md, rules, skills, skill refs) for broken paths/commands, redundant instructions, and usefulness using a stack-agnostic dual-judge protocol with planted traps. HIGH PRIORITY questionnaires at top: Q1 optional docs, Q2 B/C budget before Track A (certainty/tokens). A always runs after Q2; B/C opt-in. ADRs/RFCs excluded from T2. Mixed apply uses 11-mixed-apply.md (KEEP/CUT). Use when the user says harness eval, harness-eval, harness debug, audit AGENTS.md, audit skills/rules, instruction audit, redundancy of agent instructions, usefulness of skills, Ship/Review/Hold/Slim/Keep-core for harness, or wants Track A/B/C harness evaluation. Do NOT use for harness setup or init, feature spec-driven work (tlc-spec-driven), or applying Ship/Slim trims unless the user explicitly asks after the report." license: CC-BY-4.0 metadata: author: Tech Leads Club - github.com/tech-leads-club version: 1.8.3
Harness Eval
Run a full, stack-agnostic harness evaluation and stop at reports. Do not auto-edit AGENTS.md or skills unless the user explicitly asks after reviewing Ship/Slim.
User questionnaires (HIGH PRIORITY)
Stop and ask before continuing. Do not skip these gates. Do not silently include optional docs or spawn B/C judges.
Order after inventory: Q1 (if needed) → Q2 → then Track A (A always runs) → B/C only if approved.
Q1 — Optional project docs (after inventory)
When optional-docs-candidates.md lists optional types, ask before Q2 / Track A:
Inventory found cited project docs outside the agent skill trees.
- **Always in scope:** skill-tree files (`.agents/skills`, `.cursor/skills`, `.claude/skills`)
- **Always excluded:** ADRs / RFCs / decision-record trees (never scored as T2)
- **Optional (default: omit):** see types/paths in `optional-docs-candidates.md`
Include any optional doc types or paths in this run?
Reply with: `none` (default), type ids (e.g. `docs`), and/or specific paths.
Re-run inventory with --include-doc-type / --include-doc only after the user answers. If no optional types, skip Q1.
Q2 — Tracks B and C (before Track A — budget)
Ask before Track A so the user sets spend up front. Track A always runs next (deterministic, ~0 model tokens). B/C run only if approved.
Choose eval scope for this run (before Track A).
| Track | Question | Certainty | Token consumption |
|-------|----------|-----------|-------------------|
| **A — Correctness** | Cited path/command exists? | **Highest** — script only, no LLM. Prefers false negatives over false BROKEN. | **~0 model tokens** (always runs next) |
| **B — Redundancy** | Would an agent rediscover this cheaply without the harness? | **Medium** — dual LLM + plants; Ship only if trap PASS and both agree. Disagree → Hold. Less model-sensitive than C. | **High** — 2 judges × every claim (~N in this inventory). Each may spot-check the repo. |
| **C — Usefulness** | Does this surface change behavior vs theory/demo/overlap? | **Lowest / most subjective** — dual LLM + plants + fan-in; **model-sensitive**. Slim/Mixed need gates; prefer second-model check before large deletes. | **Highest** — 2 judges × every surface (whole files; often dominates the run). |
Notes: Ship (B) ≠ Slim (C). Rediscoverable ≠ useless. A always runs; B/C are optional.
Reply with one of: `A only`, `B`, `C`, or `B+C`.
Fill claim count from claims.md when known; surface count ≈ T0+T1+T2 markdown after extract (or say “after surfaces_extract” if not run yet).
A only: run Track A; present04; stop (no B/C judges).B: Track A, then Steps 4–6.C: Track A, then Steps 7–10 (C does not need B).B+C: Track A, then Steps 4–11.
If the user already requested B/C/full eval in the triggering message, treat as approval — still show the Q2 table once so costs are visible.
Loading this skill's files
This skill is self-contained. Protocol, scripts, and judge prompts live under this skill directory (the folder that contains this SKILL.md). Resolve SKILL_DIR as that directory — never assume another install path.
- Read references/PROTOCOL.md completely before the first run in a session (and again if scripts fail).
- Read references/judge-prompts.md when spawning Track B or Track C judges.
- Plain-language terms: references/GLOSSARY.md (also embedded at the top of
04/07/10reports). - Claim record shape: references/claims.schema.json (for tooling; agents do not need to load it every run).
- Run scripts as
python3 "$SKILL_DIR/scripts/<name>.py" ....
Run outputs (not protocol) go to the target repo at .harness-eval/runs/<run-id>/.
Critical rules
- Report-only by default. Judgment ≠ remediation.
- README out of scope as harness surface and as rediscovery/usefulness evidence.
- Stack-agnostic. Never hard-code package managers, DBs, frameworks, or folder layouts in prompts or plants. Discover manifests that exist (JS, Python, Make/Task, Rust, Go, PHP, Ruby/Rails, Java/Gradle/Maven, plus
bin/*). - Doc scope. T2 always includes agent skill-tree refs (
.agents/skills,.cursor/skills,.claude/skills). ADRs / RFCs (decision-record trees) are always excluded from T2 surfaces. Other cited project docs are optional — default omit; ask via Q1 at the top of this skill, then re-run with--include-doc-type/--include-doc. - Track A always runs after inventory (deterministic, high-precision). Prefer false negatives over false BROKEN. Placeholders (
SPEC_FOLDER,{x},[feature]) are never BROKEN. Never normalize paths withstr.lstrip('./'). - Tracks B and C require user approval via Q2 before Track A. Do not spawn B/C judges until the user opts in. User may approve B only, C only, both, or A only.
- Track B needs dual judges + plants. Judge2 is blind (must not read Judge1 scores or
trap-key.json). Ship only if trap gate PASS and dual REDUNDANT with Judge2 cost ≤ 1. - Track C needs dual judges + plants. Blind Judge2 must not read
08-usefulness-j1.mdorusefulness-trap-key.json. Slim only if trap PASS, dual SLIM/ROUTING-ONLY, and fan-in PASS (no other harness surface hard-loads the path as SoT — merge enforces this on the full skill tree, not just--seed). Usefulness is model-sensitive — recordmodel: <id>in both score files; prefer same model within a run; re-judge on a second model before large Slim deletes. - KEEP / KEEP-CORE plants must not be verbatim copies of claims/surfaces already in the deck.
- Subagents: use an allowlisted non-fast model (prefer the same family as the parent when policy allows). Do not use
*-fastmodels. - Do not equate tracks. Track B Ship ≠ Track C Slim. Rediscoverable ≠ useless; useful ≠ non-redundant.
- Slim apply / fan-in. Never stub or delete a Slim path listed under “Slim fan-in blocked” (or when
python3 "$SKILL_DIR/scripts/slim_fanin.py" --path <P>reports citers) unless those consumers are updated in the same change. - Mixed/Slim apply stays self-contained. Cutting REPO-DEMONSTRATED / THEORY means delete or compress that bulk in the harness surface. Never replace a fenced teaching snippet (or the contract it carried) with
See app/.../lib/.../test/...— that swaps SoT for a code-tree pointer. Judge evidence paths stay in score tables only; if the behavior-changing contract must survive, keep a short in-skill rule or snippet. - Mixed apply is mechanical. Dual MIXED alone is not enough. Merge emits
11-mixed-apply.mdwith per-ID KEEP (from Keep-core columns) and CUT (from Slim columns). Apply agents must follow that file only — do not re-judge, redesign, or invent a different pattern than KEEP. Empty Keep-core/Slim cells → skip that path (Hold).
Instructions
Step 1: Resolve SKILL_DIR
Set SKILL_DIR to the directory containing this SKILL.md. Verify:
$SKILL_DIR/references/PROTOCOL.md$SKILL_DIR/scripts/inventory_extract.py$SKILL_DIR/scripts/track_a_correctness.py$SKILL_DIR/scripts/merge_agreement.py$SKILL_DIR/scripts/surfaces_extract.py$SKILL_DIR/scripts/merge_usefulness.py$SKILL_DIR/scripts/slim_fanin.py$SKILL_DIR/scripts/doc_scope.py
If missing, the skill install is broken — stop.
Step 2: Inventory + claim deck
From the target repo root:
RUN_ID=$(date -u +%Y-%m-%d)-full
python3 "$SKILL_DIR/scripts/inventory_extract.py" --root . --run-id "$RUN_ID"
# Optional scope: AGENTS.md + one-hop related skills only
# python3 "$SKILL_DIR/scripts/inventory_extract.py" --root . --run-id "$RUN_ID" --seed AGENTS.md
Expected under .harness-eval/runs/$RUN_ID/: inventory.json, claims.jsonl, claims.md, trap-key.json, optional-docs-candidates.md (+ .json).
Step 2b: Optional docs — Q1 (see top)
Read optional-docs-candidates.md. If optional types exist, run Q1 from User questionnaires. Re-run inventory only after approval:
python3 "$SKILL_DIR/scripts/inventory_extract.py" --root . --run-id "$RUN_ID" \
--include-doc-type docs # and/or --include-doc path
Step 2c: Track budget — Q2 (see top)
Run Q2 from User questionnaires before Track A. Record the answer (A only / B / C / B+C). Do not start Steps 4+ unless B and/or C were approved.
Step 3: Track A (deterministic) — always run
python3 "$SKILL_DIR/scripts/track_a_correctness.py" --root . --run-id "$RUN_ID"
Expected: 04-correctness.md (includes term definitions at top). Spot-check that .agents/... cites resolve (not agents/...).
Summarize Track A (broken count + notable clusters). If Q2 was A only, stop. Otherwise continue to the approved B and/or C steps.
Step 4: Track B — Judge1
Read references/judge-prompts.md (Track B Judge1). Spawn an independent subagent with an allowlisted model. Point it at .harness-eval/runs/$RUN_ID/claims.md. It writes 05-redundancy-j1.md (include model: <id>).
Judge1 may read inventory.json. Must not read trap-key.json.
Step 5: Track B — Judge2 (blind)
Read references/judge-prompts.md (Track B Judge2). Spawn a second subagent. Writes 06-blind-scores.md.
Forbidden for Judge2: trap-key.json, 05-redundancy-j1.md, 07-agreement.md, prior agreement reports.
Prefer Steps 4 and 5 in parallel.
Step 6: Merge Track B agreement
python3 "$SKILL_DIR/scripts/merge_agreement.py" --run-dir .harness-eval/runs/$RUN_ID
Expected: 07-agreement.md (Ship/Review/Hold + What these words mean). On trap FAIL: fix plants per PROTOCOL, rescore P00x, re-merge — do not Ship.
Step 7: Track C — surface deck
python3 "$SKILL_DIR/scripts/surfaces_extract.py" --root . --run-id "$RUN_ID"
Expected: surfaces.md, surfaces.json, usefulness-trap-key.json.
Step 8: Track C — Usefulness Judge1
Read references/judge-prompts.md (Usefulness Judge1). Spawn subagent with allowlisted model (record same id in header). Writes 08-usefulness-j1.md.
Must not read usefulness-trap-key.json.
Step 9: Track C — Usefulness Judge2 (blind)
Read Usefulness Judge2 prompt. Prefer same model as Step 8 for agreement stability. Writes 09-usefulness-j2.md.
Forbidden: usefulness-trap-key.json, 08-usefulness-j1.md, 10-usefulness-agreement.md, and using Track B 05/06/07 to decide usefulness classes.
Prefer Steps 8 and 9 in parallel.
Step 10: Merge Track C agreement
python3 "$SKILL_DIR/scripts/merge_usefulness.py" --run-dir .harness-eval/runs/$RUN_ID
Expected: 10-usefulness-agreement.md (Slim/Keep-core/Mixed/Hold + What these words mean), 11-mixed-apply.md (KEEP/CUT per Mixed ID), plus slim-fanin.json. On t
Truncated for display — read the full file on GitHub.
Related Skills
ai-job-search
44.2kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.7kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
