SkillAgentSearch skills...

harness-eval

Evaluate a repo agent harness (AGENTS.md, rules, skills, skill refs) for broken paths/commands, redundant instructions, and usefulness using a stack-agnostic dual-judge protocol with planted traps.

Install / Use

npx skills add tech-leads-club/agent-skills --skill harness-eval

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

96/100

Supported Platforms

OpenAI Codex

Tags

Our assessment of harness-eval

harness-eval scores 96/100 on our quality scale, 187th of 3,044 Development & Engineering skills we index (top 7%).

Its SKILL.md is 15 KB long, well organised into 36 sections with 8 code examples: a thorough specification that gives an agent plenty to work with.

With 6,832 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
16/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 7 days ago, so harness-eval is actively maintained.
  • No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
  • Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.

Automated pattern scan on 2026-09-28. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.

harness-eval compared with similar skills

All 4 of these similar skills score higher than harness-eval; compare them before choosing.

SkillScoreStarsUpdatedFormat
harness-eval (this skill)by tech-leads-club966.8k7d agoSKILL.md
ai-job-searchby MadsLorentzen10044.2ktodayCLAUDE.md
claude-howtoby luongnv8910041.7k1d agoCLAUDE.md
algorithmic-artby anthropics100177.9k5d agoSKILL.md
pptxby anthropics100177.9k5d agoSKILL.md

Frequently asked questions

How do I install harness-eval?
Run npx skills add tech-leads-club/agent-skills --skill harness-eval. The install tabs above show the steps for each supported agent.
Which AI agents does harness-eval work with?
It is written for OpenAI Codex, as a SKILL.md file. Other agents that read the same format can often use it too.
Is harness-eval safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is harness-eval still maintained?
The repository was last updated 7 days ago, so harness-eval is actively maintained.

name: harness-eval description: "Evaluate a repo agent harness (AGENTS.md, rules, skills, skill refs) for broken paths/commands, redundant instructions, and usefulness using a stack-agnostic dual-judge protocol with planted traps. HIGH PRIORITY questionnaires at top: Q1 optional docs, Q2 B/C budget before Track A (certainty/tokens). A always runs after Q2; B/C opt-in. ADRs/RFCs excluded from T2. Mixed apply uses 11-mixed-apply.md (KEEP/CUT). Use when the user says harness eval, harness-eval, harness debug, audit AGENTS.md, audit skills/rules, instruction audit, redundancy of agent instructions, usefulness of skills, Ship/Review/Hold/Slim/Keep-core for harness, or wants Track A/B/C harness evaluation. Do NOT use for harness setup or init, feature spec-driven work (tlc-spec-driven), or applying Ship/Slim trims unless the user explicitly asks after the report." license: CC-BY-4.0 metadata: author: Tech Leads Club - github.com/tech-leads-club version: 1.8.3

Harness Eval

Run a full, stack-agnostic harness evaluation and stop at reports. Do not auto-edit AGENTS.md or skills unless the user explicitly asks after reviewing Ship/Slim.

User questionnaires (HIGH PRIORITY)

Stop and ask before continuing. Do not skip these gates. Do not silently include optional docs or spawn B/C judges.

Order after inventory: Q1 (if needed) → Q2 → then Track A (A always runs) → B/C only if approved.

Q1 — Optional project docs (after inventory)

When optional-docs-candidates.md lists optional types, ask before Q2 / Track A:

Inventory found cited project docs outside the agent skill trees.

- **Always in scope:** skill-tree files (`.agents/skills`, `.cursor/skills`, `.claude/skills`)
- **Always excluded:** ADRs / RFCs / decision-record trees (never scored as T2)
- **Optional (default: omit):** see types/paths in `optional-docs-candidates.md`

Include any optional doc types or paths in this run?
Reply with: `none` (default), type ids (e.g. `docs`), and/or specific paths.

Re-run inventory with --include-doc-type / --include-doc only after the user answers. If no optional types, skip Q1.

Q2 — Tracks B and C (before Track A — budget)

Ask before Track A so the user sets spend up front. Track A always runs next (deterministic, ~0 model tokens). B/C run only if approved.

Choose eval scope for this run (before Track A).

| Track | Question | Certainty | Token consumption |
|-------|----------|-----------|-------------------|
| **A — Correctness** | Cited path/command exists? | **Highest** — script only, no LLM. Prefers false negatives over false BROKEN. | **~0 model tokens** (always runs next) |
| **B — Redundancy** | Would an agent rediscover this cheaply without the harness? | **Medium** — dual LLM + plants; Ship only if trap PASS and both agree. Disagree → Hold. Less model-sensitive than C. | **High** — 2 judges × every claim (~N in this inventory). Each may spot-check the repo. |
| **C — Usefulness** | Does this surface change behavior vs theory/demo/overlap? | **Lowest / most subjective** — dual LLM + plants + fan-in; **model-sensitive**. Slim/Mixed need gates; prefer second-model check before large deletes. | **Highest** — 2 judges × every surface (whole files; often dominates the run). |

Notes: Ship (B) ≠ Slim (C). Rediscoverable ≠ useless. A always runs; B/C are optional.

Reply with one of: `A only`, `B`, `C`, or `B+C`.

Fill claim count from claims.md when known; surface count ≈ T0+T1+T2 markdown after extract (or say “after surfaces_extract” if not run yet).

  • A only: run Track A; present 04; stop (no B/C judges).
  • B: Track A, then Steps 4–6.
  • C: Track A, then Steps 7–10 (C does not need B).
  • B+C: Track A, then Steps 4–11.

If the user already requested B/C/full eval in the triggering message, treat as approval — still show the Q2 table once so costs are visible.

Loading this skill's files

This skill is self-contained. Protocol, scripts, and judge prompts live under this skill directory (the folder that contains this SKILL.md). Resolve SKILL_DIR as that directory — never assume another install path.

Run outputs (not protocol) go to the target repo at .harness-eval/runs/<run-id>/.

Critical rules

  1. Report-only by default. Judgment ≠ remediation.
  2. README out of scope as harness surface and as rediscovery/usefulness evidence.
  3. Stack-agnostic. Never hard-code package managers, DBs, frameworks, or folder layouts in prompts or plants. Discover manifests that exist (JS, Python, Make/Task, Rust, Go, PHP, Ruby/Rails, Java/Gradle/Maven, plus bin/*).
  4. Doc scope. T2 always includes agent skill-tree refs (.agents/skills, .cursor/skills, .claude/skills). ADRs / RFCs (decision-record trees) are always excluded from T2 surfaces. Other cited project docs are optional — default omit; ask via Q1 at the top of this skill, then re-run with --include-doc-type / --include-doc.
  5. Track A always runs after inventory (deterministic, high-precision). Prefer false negatives over false BROKEN. Placeholders (SPEC_FOLDER, {x}, [feature]) are never BROKEN. Never normalize paths with str.lstrip('./').
  6. Tracks B and C require user approval via Q2 before Track A. Do not spawn B/C judges until the user opts in. User may approve B only, C only, both, or A only.
  7. Track B needs dual judges + plants. Judge2 is blind (must not read Judge1 scores or trap-key.json). Ship only if trap gate PASS and dual REDUNDANT with Judge2 cost ≤ 1.
  8. Track C needs dual judges + plants. Blind Judge2 must not read 08-usefulness-j1.md or usefulness-trap-key.json. Slim only if trap PASS, dual SLIM/ROUTING-ONLY, and fan-in PASS (no other harness surface hard-loads the path as SoT — merge enforces this on the full skill tree, not just --seed). Usefulness is model-sensitive — record model: <id> in both score files; prefer same model within a run; re-judge on a second model before large Slim deletes.
  9. KEEP / KEEP-CORE plants must not be verbatim copies of claims/surfaces already in the deck.
  10. Subagents: use an allowlisted non-fast model (prefer the same family as the parent when policy allows). Do not use *-fast models.
  11. Do not equate tracks. Track B Ship ≠ Track C Slim. Rediscoverable ≠ useless; useful ≠ non-redundant.
  12. Slim apply / fan-in. Never stub or delete a Slim path listed under “Slim fan-in blocked” (or when python3 "$SKILL_DIR/scripts/slim_fanin.py" --path <P> reports citers) unless those consumers are updated in the same change.
  13. Mixed/Slim apply stays self-contained. Cutting REPO-DEMONSTRATED / THEORY means delete or compress that bulk in the harness surface. Never replace a fenced teaching snippet (or the contract it carried) with See app/... / lib/... / test/... — that swaps SoT for a code-tree pointer. Judge evidence paths stay in score tables only; if the behavior-changing contract must survive, keep a short in-skill rule or snippet.
  14. Mixed apply is mechanical. Dual MIXED alone is not enough. Merge emits 11-mixed-apply.md with per-ID KEEP (from Keep-core columns) and CUT (from Slim columns). Apply agents must follow that file only — do not re-judge, redesign, or invent a different pattern than KEEP. Empty Keep-core/Slim cells → skip that path (Hold).

Instructions

Step 1: Resolve SKILL_DIR

Set SKILL_DIR to the directory containing this SKILL.md. Verify:

  • $SKILL_DIR/references/PROTOCOL.md
  • $SKILL_DIR/scripts/inventory_extract.py
  • $SKILL_DIR/scripts/track_a_correctness.py
  • $SKILL_DIR/scripts/merge_agreement.py
  • $SKILL_DIR/scripts/surfaces_extract.py
  • $SKILL_DIR/scripts/merge_usefulness.py
  • $SKILL_DIR/scripts/slim_fanin.py
  • $SKILL_DIR/scripts/doc_scope.py

If missing, the skill install is broken — stop.

Step 2: Inventory + claim deck

From the target repo root:

RUN_ID=$(date -u +%Y-%m-%d)-full
python3 "$SKILL_DIR/scripts/inventory_extract.py" --root . --run-id "$RUN_ID"
# Optional scope: AGENTS.md + one-hop related skills only
# python3 "$SKILL_DIR/scripts/inventory_extract.py" --root . --run-id "$RUN_ID" --seed AGENTS.md

Expected under .harness-eval/runs/$RUN_ID/: inventory.json, claims.jsonl, claims.md, trap-key.json, optional-docs-candidates.md (+ .json).

Step 2b: Optional docs — Q1 (see top)

Read optional-docs-candidates.md. If optional types exist, run Q1 from User questionnaires. Re-run inventory only after approval:

python3 "$SKILL_DIR/scripts/inventory_extract.py" --root . --run-id "$RUN_ID" \
  --include-doc-type docs   # and/or --include-doc path

Step 2c: Track budget — Q2 (see top)

Run Q2 from User questionnaires before Track A. Record the answer (A only / B / C / B+C). Do not start Steps 4+ unless B and/or C were approved.

Step 3: Track A (deterministic) — always run

python3 "$SKILL_DIR/scripts/track_a_correctness.py" --root . --run-id "$RUN_ID"

Expected: 04-correctness.md (includes term definitions at top). Spot-check that .agents/... cites resolve (not agents/...).

Summarize Track A (broken count + notable clusters). If Q2 was A only, stop. Otherwise continue to the approved B and/or C steps.

Step 4: Track B — Judge1

Read references/judge-prompts.md (Track B Judge1). Spawn an independent subagent with an allowlisted model. Point it at .harness-eval/runs/$RUN_ID/claims.md. It writes 05-redundancy-j1.md (include model: <id>).

Judge1 may read inventory.json. Must not read trap-key.json.

Step 5: Track B — Judge2 (blind)

Read references/judge-prompts.md (Track B Judge2). Spawn a second subagent. Writes 06-blind-scores.md.

Forbidden for Judge2: trap-key.json, 05-redundancy-j1.md, 07-agreement.md, prior agreement reports.

Prefer Steps 4 and 5 in parallel.

Step 6: Merge Track B agreement

python3 "$SKILL_DIR/scripts/merge_agreement.py" --run-dir .harness-eval/runs/$RUN_ID

Expected: 07-agreement.md (Ship/Review/Hold + What these words mean). On trap FAIL: fix plants per PROTOCOL, rescore P00x, re-merge — do not Ship.

Step 7: Track C — surface deck

python3 "$SKILL_DIR/scripts/surfaces_extract.py" --root . --run-id "$RUN_ID"

Expected: surfaces.md, surfaces.json, usefulness-trap-key.json.

Step 8: Track C — Usefulness Judge1

Read references/judge-prompts.md (Usefulness Judge1). Spawn subagent with allowlisted model (record same id in header). Writes 08-usefulness-j1.md.

Must not read usefulness-trap-key.json.

Step 9: Track C — Usefulness Judge2 (blind)

Read Usefulness Judge2 prompt. Prefer same model as Step 8 for agreement stability. Writes 09-usefulness-j2.md.

Forbidden: usefulness-trap-key.json, 08-usefulness-j1.md, 10-usefulness-agreement.md, and using Track B 05/06/07 to decide usefulness classes.

Prefer Steps 8 and 9 in parallel.

Step 10: Merge Track C agreement

python3 "$SKILL_DIR/scripts/merge_usefulness.py" --run-dir .harness-eval/runs/$RUN_ID

Expected: 10-usefulness-agreement.md (Slim/Keep-core/Mixed/Hold + What these words mean), 11-mixed-apply.md (KEEP/CUT per Mixed ID), plus slim-fanin.json. On t

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars6.8k
CategoryDevelopment
Updated7d ago
Forks551

Languages

TypeScript

Trust signals

88/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 medium