spec-driven-eval
Scores how completely an implementation fulfills a PRD/spec, case by case, and produces a single comparable final grade. Invoke only when explicitly named (e.g. run spec-driven-eval); do not auto-trigger
Install / Use
npx skills add tech-leads-club/agent-skills --skill spec-driven-evalInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Content & MediaSupported Platforms
Tags
Our assessment of spec-driven-eval
spec-driven-eval scores 96/100 on our quality scale, 50th of 710 Content & Media skills we index (top 8%).
Its SKILL.md is 34 KB long, well organised into 23 sections with 9 code examples: a thorough specification that gives an agent plenty to work with.
With 6,832 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 7 days ago, so spec-driven-eval is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-09-28. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
spec-driven-eval compared with similar skills
All 4 of these similar skills score higher than spec-driven-eval; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| spec-driven-eval (this skill)by tech-leads-club | 96 | 6.8k | 7d ago | SKILL.md |
| siyuanby siyuan-note | 100 | 46.5k | today | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 5d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 5d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 6d ago | SKILL.md |
Frequently asked questions
- How do I install spec-driven-eval?
- Run
npx skills add tech-leads-club/agent-skills --skill spec-driven-eval. The install tabs above show the steps for each supported agent. - Which AI agents does spec-driven-eval work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is spec-driven-eval safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is spec-driven-eval still maintained?
- The repository was last updated 7 days ago, so spec-driven-eval is actively maintained.
Skill content
View source on GitHubname: spec-driven-eval description: Scores how completely an implementation fulfills a PRD/spec, case by case, and produces a single comparable final grade. Invoke only when explicitly named (e.g. run spec-driven-eval); do not auto-trigger. Use when benchmarking spec-driven implementations, grading acceptance criteria, evaluating whether a feature was 100% implemented, comparing multiple implementations of the same PRD, or auditing implementation and test coverage (unit and e2e) against product requirements. Do NOT use for planning or building features (use tlc-spec-driven), writing PRDs, or general code review unrelated to a spec. license: CC-BY-4.0 disable-model-invocation: true metadata: author: Waldemar Neto - github.com/waldemarnt version: '1.0.0'
Spec-Driven Implementation Evaluation
Evaluate a spec-driven-development (SDD) effort against a PRD case by case (acceptance criterion by acceptance criterion), scoring implementation and tests separately, and roll up to a single comparable final grade. Designed for benchmarking: the same PRD evaluated across different SDD frameworks — and the same effort evaluated twice — must yield comparable, reproducible numbers.
Two subjects, two questions
This eval answers two independent questions, and keeps their verdicts separate because they fail independently:
- How good is the framework at respecting and extracting requirements? — Did it honor the PRD and stay in bounds (respect:
Finalimplementation side + ScopeS), and did it surface the implicit requirements the PRD only implied, without noise (extract: ElicitationE)? - How good is the harness at ensuring they are all implemented? — Does the test suite prove every sanctioned requirement is actually built (
T+ Engineering GatesG)?
The linkage that makes "all implemented" well-defined: the harness is held accountable for the full sanctioned requirement set =
PRD acceptance criteria ∪ valid E-additions. A valid requirement the framework extracted but the harness never tests is a harness miss (T), not a framework miss (Ekeeps the extraction credit). Extraction defines the verification target.
Scoring is checklist-based: every criterion is decomposed into atomic binary (MET / UNMET) checks, each backed by file:line evidence. Binary decomposition is the design choice that makes the grade reproducible — graded/Likert scales (is this a 3 or a 4?) are the dominant source of evaluator-to-evaluator disagreement; binary checks raise inter-evaluator agreement to roughly human level. Partial credit is derived from the fraction of checks met, never judged on a sliding scale.
When to use
- "Evaluate/score this PRD case by case and give a final grade"
- "Was this story/feature implemented 100%?"
- Benchmarking multiple spec-driven implementations of the same PRD
- Auditing implementation and test coverage (unit + e2e) against acceptance criteria
Inputs required
- The PRD (ground-truth product intent) — the user story acceptance criteria are the unit of evaluation.
- The implementation (production code).
- The tests (unit + e2e).
- The SDD-derived artifacts —
spec.mdandtasks.md(refined ACs, requirement IDs, derived requirements). Required for the ElicitationEand ScopeSaxes (they grade these against the PRD). When absent,Final/Tstill run, but reportE/Sasn/a — no derived spec. The PRD remains the source of truth for what counts as "expected"; the derived spec is what gets graded for respect and extraction.
Scoping the diff. Before scoring, use git to identify which files changed for this implementation. That diff surface is the primary search scope for all file:line evidence in steps 4 and 8:
git diff <base>..<head> --name-only # branch or PR
git diff --name-only HEAD # uncommitted changes
git status --short # include untracked new files
Record the diff surface in the report. Evidence outside it is still valid (e.g. a pre-existing file was modified), but note when a check relies on files not in the diff — that may indicate the wrong base was chosen.
Quick start
New to running this end-to-end? See quickstart.md for the 4-chat-session flow (freeze baseline → plan → implement → evaluate) with paste-ready prompts. Read it when you need the operational how-to; the rest of this file is the scoring methodology.
Core rules (read first)
These five rules govern every score. They exist to make the grade auditable and reproducible.
- Evidence or zero. Every MET check MUST cite evidence as
file:line(orfile:startLine-endLine). No located evidence ⇒ the check is UNMET. Never award credit from assumptions or from the PRD restating intent. - Search-before-zero (anti false-negative). Before marking a check UNMET for "not found", record the search actually performed — starting with the diff-surface files from the scoping step, then the grep/glob terms tried and any additional files/dirs inspected. A check scored UNMET must confirm the behavior is absent from the diff surface, not just from an ad-hoc grep across the full repo. UNMET means searched and genuinely absent, not did not look. If no search is shown, the check is not yet scored, not UNMET.
- Read the path end-to-end — including the data shape. A check is MET only if you traced the real code/test path, not because a symbol name matches. For an emitted/returned/persisted artifact, inspect the constructed payload object itself, not just the call site — a present
emit(...)/return ...does not prove the named field is in the payload (see the Conjunction rule). A test that exercises a behavior but does not assert it does not meet a verification check. - Judge ≠ author for benchmarks. When grading to compare implementations, the evaluating model should differ from the model that wrote the code; LLM judges over-reward their own output (self-preference bias). If they are the same, flag it under Assumptions and treat borderline checks as UNMET.
- The evaluator is read-only over the subject. Never modify the code under evaluation — not to fix a failing gate, not to "help", not even for trivial type errors. A red gate stays red: record
✗, apply theAdjusted Final, and put the required correction in the report's fix list. Modifying the subject mid-evaluation contaminates the benchmark (the grade no longer measures the framework's output) and invalidates the diff surface. If a correction was accidentally applied, revert it and score the original state.
Reproducibility rules
The grade must be stable: the same PRD + same commit must yield the same Final across runs and evaluators. Obey all of:
- Freeze the AC list AND its checklist (the single biggest drift source). AC enumeration and check granularity are decided once and pinned, because both
I,T, and theStory_scoredenominator depend on them:- If
spec.mdwith stable requirement IDs exists, the AC list and IDs are frozen — use them verbatim; do not re-split or merge. - Derive the binary checklist for each AC once and persist it to
<spec-folder>/evaluations/_ac-baseline.md(IDs + verbatim AC text + the I-checks and T-checks). Every implementation of that PRD is scored against the identical checklist. - One AC = one testable assertion of intent; one check = one atomic, observable yes/no proposition. Do not collapse distinct behaviors into one check, nor split one behavior across several.
- If
- Priority comes from the PRD, never inferred silently. Use the PRD's explicit P0/P1/P2 labels. If a story is unlabeled, mark its priority
ASSUMEDin the report and list it under Assumptions; never let an inferred priority silently move the3/2/0weights. - Round once, at the end. Carry full precision through the arithmetic; round only the reported
AC_score,Story_score, andFinalto 2 decimals. Band assignment uses the roundedFinal. - Compute, don't calculate. The roll-up (
Σw, everyStory_score,Final,Adjusted Final) MUST be computed by executing a script (e.g.node -e/python3 -c) that takes the per-ACI/Tfractions and the priority-weight table as input, with the script's output pasted into the report. Never do the arithmetic mentally or by hand — a hand-summed denominator (Σw) is a known failure mode that silently shifts the grade band.Σwmust be derived inside the script from the same priority table shown in the report, not typed in as a literal. - Self-consistency ensemble (k = 3). Evaluate the checklist three times independently at low/zero temperature and take the majority MET/UNMET per check before computing any number. Majority voting over k=3 removes most stochastic flips at low cost. If the three passes disagree on a check, that check is borderline — keep the majority verdict and note it; persistent disagreement means the check wording is ambiguous (sharpen it in the baseline, see Calibration).
Scoring model
Unit of scoring: the acceptance criterion (AC), decomposed into binary checks
For each in-scope AC, the frozen checklist holds two sets of atomic checks, each scored MET (1) / UNMET (0):
Implementation checks (I-checks) — one per distinct observable behavior the AC asserts. A check is MET only if production code performs that behavior, evidenced file:line. Include only behaviorally-observable clauses (a verb the AC states: creates / returns / rejects / persists / grants / defaults-to). Non-functional polish (wording, logging) is not a check — note it separately so it never moves the score.
When parsing an AC into checks (done once, at baseline-freeze time), apply both rules below so verbs aren't the only thing captured:
Conjunction / payload-field rule. When an AC enumerates multiple items — joined by "and", ",", "with [field]", "including" — especially in the subject or payload of an emitted/returned/persisted artifact, treat each named field or entity as its own I-check. Do not stop at confirming the parent method is called: for each named field, open the actual data object constructed at file:line and verify the field is present in it (score against the payload shape, not the call site).
Example: "emit a trigger associated with the user and the trial end date" → two I-checks: (1) trigger payload carries userId, (2) trigger payload carries trialEndsAt. Confirming emit(...) is reached does not satisfy check (2) if the payload object omits the field.
Disjunction / product-chosen rule. When an AC presents alternatives ("A or B", "pause or cancel", "immediately or at period end"), decide the reading once and freeze it in the baseline (the reading itself must not vary per run):
- Independent paths — both are separately reachable behaviors (e.g. "cancel immediately or at period end"): one I-check per path, each independently MET/UNMET.
- Product-controlled / configurable — the AC frames the choice as product-driven (e.g. "apply the product-chosen behavior: pause or cancel"): two I-checks — (1) the recommended/default behavior is implemented, and (2) the alternative is reachable without a code change (config key / flag / env). Hard-coding one option with no switch to the other ⇒ check (2) UNMET. A different feature that happens to reach the other option (e.g. user-initiated cancel) does not satisfy (2) — it must be the same product-controlled decision point.
- If "product-chosen" is itself ambiguous between runtime-configurable and design-time chosen, resolve it against the PRD/spec wording and record the resolved reading in the baseline; only require reachability of the non-default path when the AC frames the choice as runtime/product-controlled.
**Wiring / in
Truncated for display — read the full file on GitHub.
Related Skills
siyuan
46.5kAn open-source, privacy-first, self-hosted knowledge workspace where humans and AI agents work together 开源、隐私优先、自托管的知识工作空间,让人与智能体在此协作
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
