SkillAgentSearch skills...

spec-driven-eval

Scores how completely an implementation fulfills a PRD/spec, case by case, and produces a single comparable final grade. Invoke only when explicitly named (e.g. run spec-driven-eval); do not auto-trigger

Install / Use

npx skills add tech-leads-club/agent-skills --skill spec-driven-eval

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

96/100

Supported Platforms

Universal

Tags

Our assessment of spec-driven-eval

spec-driven-eval scores 96/100 on our quality scale, 50th of 710 Content & Media skills we index (top 8%).

Its SKILL.md is 34 KB long, well organised into 23 sections with 9 code examples: a thorough specification that gives an agent plenty to work with.

With 6,832 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
16/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 7 days ago, so spec-driven-eval is actively maintained.
  • No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
  • Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.

Automated pattern scan on 2026-09-28. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.

spec-driven-eval compared with similar skills

All 4 of these similar skills score higher than spec-driven-eval; compare them before choosing.

SkillScoreStarsUpdatedFormat
spec-driven-eval (this skill)by tech-leads-club966.8k7d agoSKILL.md
siyuanby siyuan-note10046.5ktodayMCP Server
algorithmic-artby anthropics100177.9k5d agoSKILL.md
pptxby anthropics100177.9k5d agoSKILL.md
designby nextlevelbuilder100130.2k6d agoSKILL.md

Frequently asked questions

How do I install spec-driven-eval?
Run npx skills add tech-leads-club/agent-skills --skill spec-driven-eval. The install tabs above show the steps for each supported agent.
Which AI agents does spec-driven-eval work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is spec-driven-eval safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is spec-driven-eval still maintained?
The repository was last updated 7 days ago, so spec-driven-eval is actively maintained.

name: spec-driven-eval description: Scores how completely an implementation fulfills a PRD/spec, case by case, and produces a single comparable final grade. Invoke only when explicitly named (e.g. run spec-driven-eval); do not auto-trigger. Use when benchmarking spec-driven implementations, grading acceptance criteria, evaluating whether a feature was 100% implemented, comparing multiple implementations of the same PRD, or auditing implementation and test coverage (unit and e2e) against product requirements. Do NOT use for planning or building features (use tlc-spec-driven), writing PRDs, or general code review unrelated to a spec. license: CC-BY-4.0 disable-model-invocation: true metadata: author: Waldemar Neto - github.com/waldemarnt version: '1.0.0'

Spec-Driven Implementation Evaluation

Evaluate a spec-driven-development (SDD) effort against a PRD case by case (acceptance criterion by acceptance criterion), scoring implementation and tests separately, and roll up to a single comparable final grade. Designed for benchmarking: the same PRD evaluated across different SDD frameworks — and the same effort evaluated twice — must yield comparable, reproducible numbers.

Two subjects, two questions

This eval answers two independent questions, and keeps their verdicts separate because they fail independently:

  1. How good is the framework at respecting and extracting requirements? — Did it honor the PRD and stay in bounds (respect: Final implementation side + Scope S), and did it surface the implicit requirements the PRD only implied, without noise (extract: Elicitation E)?
  2. How good is the harness at ensuring they are all implemented? — Does the test suite prove every sanctioned requirement is actually built (T + Engineering Gates G)?

The linkage that makes "all implemented" well-defined: the harness is held accountable for the full sanctioned requirement set = PRD acceptance criteria ∪ valid E-additions. A valid requirement the framework extracted but the harness never tests is a harness miss (T), not a framework miss (E keeps the extraction credit). Extraction defines the verification target.

Scoring is checklist-based: every criterion is decomposed into atomic binary (MET / UNMET) checks, each backed by file:line evidence. Binary decomposition is the design choice that makes the grade reproducible — graded/Likert scales (is this a 3 or a 4?) are the dominant source of evaluator-to-evaluator disagreement; binary checks raise inter-evaluator agreement to roughly human level. Partial credit is derived from the fraction of checks met, never judged on a sliding scale.

When to use

  • "Evaluate/score this PRD case by case and give a final grade"
  • "Was this story/feature implemented 100%?"
  • Benchmarking multiple spec-driven implementations of the same PRD
  • Auditing implementation and test coverage (unit + e2e) against acceptance criteria

Inputs required

  1. The PRD (ground-truth product intent) — the user story acceptance criteria are the unit of evaluation.
  2. The implementation (production code).
  3. The tests (unit + e2e).
  4. The SDD-derived artifacts — spec.md and tasks.md (refined ACs, requirement IDs, derived requirements). Required for the Elicitation E and Scope S axes (they grade these against the PRD). When absent, Final/T still run, but report E/S as n/a — no derived spec. The PRD remains the source of truth for what counts as "expected"; the derived spec is what gets graded for respect and extraction.

Scoping the diff. Before scoring, use git to identify which files changed for this implementation. That diff surface is the primary search scope for all file:line evidence in steps 4 and 8:

git diff <base>..<head> --name-only   # branch or PR
git diff --name-only HEAD             # uncommitted changes
git status --short                    # include untracked new files

Record the diff surface in the report. Evidence outside it is still valid (e.g. a pre-existing file was modified), but note when a check relies on files not in the diff — that may indicate the wrong base was chosen.


Quick start

New to running this end-to-end? See quickstart.md for the 4-chat-session flow (freeze baseline → plan → implement → evaluate) with paste-ready prompts. Read it when you need the operational how-to; the rest of this file is the scoring methodology.


Core rules (read first)

These five rules govern every score. They exist to make the grade auditable and reproducible.

  1. Evidence or zero. Every MET check MUST cite evidence as file:line (or file:startLine-endLine). No located evidence ⇒ the check is UNMET. Never award credit from assumptions or from the PRD restating intent.
  2. Search-before-zero (anti false-negative). Before marking a check UNMET for "not found", record the search actually performed — starting with the diff-surface files from the scoping step, then the grep/glob terms tried and any additional files/dirs inspected. A check scored UNMET must confirm the behavior is absent from the diff surface, not just from an ad-hoc grep across the full repo. UNMET means searched and genuinely absent, not did not look. If no search is shown, the check is not yet scored, not UNMET.
  3. Read the path end-to-end — including the data shape. A check is MET only if you traced the real code/test path, not because a symbol name matches. For an emitted/returned/persisted artifact, inspect the constructed payload object itself, not just the call site — a present emit(...)/return ... does not prove the named field is in the payload (see the Conjunction rule). A test that exercises a behavior but does not assert it does not meet a verification check.
  4. Judge ≠ author for benchmarks. When grading to compare implementations, the evaluating model should differ from the model that wrote the code; LLM judges over-reward their own output (self-preference bias). If they are the same, flag it under Assumptions and treat borderline checks as UNMET.
  5. The evaluator is read-only over the subject. Never modify the code under evaluation — not to fix a failing gate, not to "help", not even for trivial type errors. A red gate stays red: record ✗, apply the Adjusted Final, and put the required correction in the report's fix list. Modifying the subject mid-evaluation contaminates the benchmark (the grade no longer measures the framework's output) and invalidates the diff surface. If a correction was accidentally applied, revert it and score the original state.

Reproducibility rules

The grade must be stable: the same PRD + same commit must yield the same Final across runs and evaluators. Obey all of:

  1. Freeze the AC list AND its checklist (the single biggest drift source). AC enumeration and check granularity are decided once and pinned, because both I, T, and the Story_score denominator depend on them:
    • If spec.md with stable requirement IDs exists, the AC list and IDs are frozen — use them verbatim; do not re-split or merge.
    • Derive the binary checklist for each AC once and persist it to <spec-folder>/evaluations/_ac-baseline.md (IDs + verbatim AC text + the I-checks and T-checks). Every implementation of that PRD is scored against the identical checklist.
    • One AC = one testable assertion of intent; one check = one atomic, observable yes/no proposition. Do not collapse distinct behaviors into one check, nor split one behavior across several.
  2. Priority comes from the PRD, never inferred silently. Use the PRD's explicit P0/P1/P2 labels. If a story is unlabeled, mark its priority ASSUMED in the report and list it under Assumptions; never let an inferred priority silently move the 3/2/0 weights.
  3. Round once, at the end. Carry full precision through the arithmetic; round only the reported AC_score, Story_score, and Final to 2 decimals. Band assignment uses the rounded Final.
  4. Compute, don't calculate. The roll-up (Σw, every Story_score, Final, Adjusted Final) MUST be computed by executing a script (e.g. node -e / python3 -c) that takes the per-AC I/T fractions and the priority-weight table as input, with the script's output pasted into the report. Never do the arithmetic mentally or by hand — a hand-summed denominator (Σw) is a known failure mode that silently shifts the grade band. Σw must be derived inside the script from the same priority table shown in the report, not typed in as a literal.
  5. Self-consistency ensemble (k = 3). Evaluate the checklist three times independently at low/zero temperature and take the majority MET/UNMET per check before computing any number. Majority voting over k=3 removes most stochastic flips at low cost. If the three passes disagree on a check, that check is borderline — keep the majority verdict and note it; persistent disagreement means the check wording is ambiguous (sharpen it in the baseline, see Calibration).

Scoring model

Unit of scoring: the acceptance criterion (AC), decomposed into binary checks

For each in-scope AC, the frozen checklist holds two sets of atomic checks, each scored MET (1) / UNMET (0):

Implementation checks (I-checks) — one per distinct observable behavior the AC asserts. A check is MET only if production code performs that behavior, evidenced file:line. Include only behaviorally-observable clauses (a verb the AC states: creates / returns / rejects / persists / grants / defaults-to). Non-functional polish (wording, logging) is not a check — note it separately so it never moves the score.

When parsing an AC into checks (done once, at baseline-freeze time), apply both rules below so verbs aren't the only thing captured:

Conjunction / payload-field rule. When an AC enumerates multiple items — joined by "and", ",", "with [field]", "including" — especially in the subject or payload of an emitted/returned/persisted artifact, treat each named field or entity as its own I-check. Do not stop at confirming the parent method is called: for each named field, open the actual data object constructed at file:line and verify the field is present in it (score against the payload shape, not the call site). Example: "emit a trigger associated with the user and the trial end date" → two I-checks: (1) trigger payload carries userId, (2) trigger payload carries trialEndsAt. Confirming emit(...) is reached does not satisfy check (2) if the payload object omits the field.

Disjunction / product-chosen rule. When an AC presents alternatives ("A or B", "pause or cancel", "immediately or at period end"), decide the reading once and freeze it in the baseline (the reading itself must not vary per run):

  • Independent paths — both are separately reachable behaviors (e.g. "cancel immediately or at period end"): one I-check per path, each independently MET/UNMET.
  • Product-controlled / configurable — the AC frames the choice as product-driven (e.g. "apply the product-chosen behavior: pause or cancel"): two I-checks — (1) the recommended/default behavior is implemented, and (2) the alternative is reachable without a code change (config key / flag / env). Hard-coding one option with no switch to the other ⇒ check (2) UNMET. A different feature that happens to reach the other option (e.g. user-initiated cancel) does not satisfy (2) — it must be the same product-controlled decision point.
  • If "product-chosen" is itself ambiguous between runtime-configurable and design-time chosen, resolve it against the PRD/spec wording and record the resolved reading in the baseline; only require reachability of the non-default path when the AC frames the choice as runtime/product-controlled.

**Wiring / in

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars6.8k
CategoryContent
Updated7d ago
Forks551

Languages

TypeScript

Trust signals

88/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 medium