measure-okr-grader
Scores completed OKR sets at cycle close with KR-level scoring per the canonical OKR type enum (committed | aspirational | learning | operational_health | compliance_or_safety), committed-vs-aspirational interpretation, evidence quality assessment, learning synthesis, and next-cycle recommendations.…
Install / Use
npx skills add product-on-purpose/pm-skills --skill measure-okr-graderInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Education & ResearchSupported Platforms
Tags
Our assessment of measure-okr-grader
measure-okr-grader scores 85/100 on our quality scale, 292nd of 426 Education & Research skills we index.
Its SKILL.md is 16 KB long, well organised into 10 sections and no code examples: a thorough specification that gives an agent plenty to work with.
It has 697 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 16 days ago, so measure-okr-grader is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
measure-okr-grader compared with similar skills
All 4 of these similar skills score higher than measure-okr-grader; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| measure-okr-grader (this skill)by product-on-purpose | 85 | 697 | 16d ago | SKILL.md |
| last30days-skillby mvanhorn | 100 | 63.5k | 3d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 13d ago | SKILL.md |
Frequently asked questions
- How do I install measure-okr-grader?
- Run
npx skills add product-on-purpose/pm-skills --skill measure-okr-grader. The install tabs above show the steps for each supported agent. - Which AI agents does measure-okr-grader work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is measure-okr-grader safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is measure-okr-grader still maintained?
- The repository was last updated 16 days ago, so measure-okr-grader is actively maintained.
Skill content
View source on GitHubname: measure-okr-grader description: Scores completed OKR sets at cycle close with KR-level scoring per the canonical OKR type enum (committed | aspirational | learning | operational_health | compliance_or_safety), committed-vs-aspirational interpretation, evidence quality assessment, learning synthesis, and next-cycle recommendations. Refuses to retroactively change targets or shrink committed scope, average away guardrail KRs, treat 0.7 as success for committed or compliance_or_safety KRs, equate effort with impact, or use scores for individual performance. Hands off to iterate-lessons-log, iterate-retrospective, define-hypothesis, measure-dashboard-requirements, measure-instrumentation-spec, and foundation-okr-writer. license: Apache-2.0 metadata: phase: measure version: "1.0.1" updated: 2026-07-04 category: reflection frameworks: [triple-diamond, okrs, lean-startup] author: product-on-purpose
<!-- PM-Skills | https://github.com/product-on-purpose/pm-skills | Apache 2.0 -->OKR Grader
An OKR Cycle Review is a backward-looking artifact that closes the loop on a completed OKR set. It scores each KR against its baseline and target, separates committed from aspirational interpretation, surfaces what evidence does and does not support, names what the team learned, and prepares input for next-cycle drafting. Done well, a cycle review protects the integrity of the OKR operating system by refusing to dress up missed commitments as aspirational stretch, refusing to celebrate effort over outcome, and refusing to let scoring carry weight it cannot bear.
This skill is an evidence interpreter, not an arithmetic engine. Its job is to read final KR values, compare them against the original OKR set's intent, and produce a review that names the learning honestly. It enforces the empirical scoring conventions drawn from Doerr (Measure What Matters), Wodtke (Radical Focus), Castro (committed vs aspirational interpretation), Grove (High Output Management), and the OKR community's accumulated practice on misuse failure modes. It pairs with foundation-okr-writer (which produced the OKR set being scored) and hands off the learnings produced here to the iterate skills that consume them.
When to Use
- The OKR cycle has ended (or you are scoring a partial-cycle close)
- You have final or interim KR values, baselines, and targets
- Stakeholders need a clear review with score, evidence, and learning
- The team is deciding what to continue, stop, change, or carry forward
- There is disagreement about whether a score is good or bad
- Evidence quality across KRs is uneven and needs to be made visible
When NOT to Use
- You are still drafting OKRs - use
foundation-okr-writer - You want a generic team retro - use
iterate-retrospective - You are reporting a single experiment result - use
measure-experiment-results - You need a stakeholder progress update without scoring - use
foundation-stakeholder-update - The OKR set was never agreed on or never tracked - scoring requires an authored set; backfill via
foundation-okr-writerfirst - You want to use scores to evaluate individuals - the skill refuses this
Instructions
When asked to score completed OKRs, follow these steps:
-
Validate scoring readiness Check inputs: original OKR set, cycle dates, final KR values (or interim values for partial-close), baselines, targets, evidence sources, and OKR types (committed | aspirational | learning | operational_health | compliance_or_safety). If a value is missing, mark it explicitly (
not-yet-observable,not-instrumented,not-supplied); never fabricate. Refuse to grade KRs whose original definitions are missing entirely. -
Classify each KR's type and indicator class The OKR type is one of
committed | aspirational | learning | operational_health | compliance_or_safety(the five values produced byfoundation-okr-writer). The indicator class is one ofleading | lagging | guardrail | health | evidence_generation. Carry both forward from the original OKR set, or assign defaults if the original set did not specify. The OKR type determines the scoring convention:aspirationaluses the 0.6 to 0.7 sweet spot;committedtargets 1.0;compliance_or_safetyis binary;operational_healthis pass | fail | drift-within-tolerance against a threshold band;learninggrades by validated or invalidated rather than by score. The indicator class adds independent rules that apply on top of the type's scoring (see Step 3). -
Score each KR Score each KR using the convention for its OKR type, then apply the indicator-class rules on top; see the Scoring Rules section below for the full per-type convention table and the guardrail rule (do not restate them here). For each score, state the calculation or rationale and the evidence confidence (high | medium | low | unknown).
-
Interpret the objective score Avoid naive averaging when one KR is a guardrail, compliance threshold, or learning KR. Produce a qualitative read of the objective alongside any rough numeric average. State explicitly what the score does and does not mean.
-
Assess evidence quality For each KR, name the evidence's reliability and any caveats (instrumentation gaps, target shifts mid-cycle, cohort definition changes, measurement window mismatches, sample-size limitations). Recommend fixes for next cycle's measurement plan.
-
Review initiatives as bets For each initiative the team ran, name which KR it was expected to move, whether it shipped, what its apparent contribution was, and whether the evidence supports continuing, retiring, or reworking it. Use Castro's "initiatives are bets, not commitments" framing. Separate ship-status from KR-impact; an initiative that shipped on time but did not move its KR is not a partial win.
-
Synthesize learning Capture validated assumptions, invalidated assumptions, surprises, and decision implications. Distinguish between learnings about the customer or product (carry forward), learnings about team process (hand to
iterate-retrospective), and learnings about measurement (hand tomeasure-instrumentation-specormeasure-dashboard-requirements). -
Prepare next-cycle recommendations For each objective: continue, revise, retire, or escalate. Suggest candidate next-cycle OKRs or open questions for
foundation-okr-writer. Hand-off measurement gaps tomeasure-dashboard-requirementsormeasure-instrumentation-spec. Hand-off assumption tests todefine-hypothesis. Hand-off team-process work toiterate-retrospective. Hand-off organizational memory toiterate-lessons-log. Hand-off next-cycle drafting tofoundation-okr-writer. -
Surface risks in interpretation Make explicit any places the score could mislead a reader: forced numeric scores on KRs that are not yet observable, confounded initiative results, stakeholder framings that under-state evidence, single-cycle results that need a second cycle of confirmation.
-
Note the source of truth The artifact is a review document, not the canonical OKR system. Include a
source_of_truthfield pointing to the original OKR tracker. -
Finalize for direct use Remove all skill instruction commentary from the final artifact. The final output should be reader-facing.
Constraint Rules (MUST / MUST NOT)
These rules are non-negotiable. The skill enforces them in every grading run.
- MUST NOT retroactively change baselines, targets, or KR definitions. If the team adjusted these mid-cycle, document the change explicitly and grade against both the original and adjusted versions.
- MUST NOT retroactively shrink the scope of a
committedorcompliance_or_safetyKR to mark partial coverage as a pass. If the original commitment named 3 healthcare accounts and only 1 has been audited, the KR isnot-yet-fully-observable. The 1-account result is a sub-signal, not the KR score. - MUST NOT treat 0.7 as success for
committed,compliance_or_safety, oroperational_healthKRs. Those target 1.0 (or the threshold band). - MUST NOT average away a failed guardrail. A failed guardrail is a separate signal that does not get diluted by the primary KR's success.
- MUST NOT equate effort with impact. Initiatives that shipped on time but failed to move their KR are not partial wins.
- MUST NOT use OKR scores as individual performance ratings or compensation inputs. If the user requests this, refuse and explain the sandbagging and learning-suppression risks.
- MUST NOT punish honest stretch when aspirational intent was explicit and disclosed at OKR-writing time. A 0.6 aspirational score is the designed sweet spot.
- MUST NOT celebrate missed committed goals as ambitious failure. Committed misses are misses.
- MUST mark any not-yet-observable KR explicitly (e.g., a 90-day retention cohort whose window extends past cycle close). Forced numeric scores on not-yet-observable KRs are misleading.
- MUST include evidence confidence on every KR score (high | medium | low | unknown).
- MUST NOT become the canonical source of truth. Always include a
source_of_truthpointer to the user's actual OKR tracker.
Scoring Rules
The skill applies these conventions to every cycle review. The convention follows the OKR type, not the team's preference at grading time. OKR type and indicator class are independent dimensions; type controls scoring, indicator class adds reporting rules.
OKR types determine the scoring convention:
aspirational: numeric score on a 0 to 1 scale = (actual - baseline) / (target - baseline). Sweet spot is 0.6 to 0.7. Below 0.4 is a miss; above 0.8 over multiple cycles suggests sandbagged targets needing recalibration.committed: pass or fail against the target. Anything below 1.0 is a miss requiring postmortem. Do not soften with aspirational interpretation.compliance_or_safety: binary. Met or not met. No partial credit. No retroactive scope shrinkage. If the committed scope is only partially observable (some audits pending, some accounts deferred), mark the KR asnot-yet-fully-observable; the observed subset is a sub-signal, not the KR score.operational_health: pass | fail | drift-within-tolerance against the threshold band.learning: validated | invalidated | partially-validated | insufficient-evidence. No numeric score.
Indicator class rules apply on top of the OKR type's scoring:
- indicator class
guardrail: the KR is scored per its OKR type, and additionally is reported as its own signal, never averaged into the primary objective score. A failed guardrail does not dilute a high primary KR score, regardless of whether the guardrail itself iscommitted,aspirational,operational_health, orcompliance_or_safety.
Special states:
not-yet-observable: score deferred. Do not force a numeric score; mark interim signal and projected score with explicit confidence and the date the final score becomes available.not-yet-fully-observable: acommittedorcompliance_or_safetyKR with partial coverage. Score the KR as deferred until full coverage is observable. Do NOT promote a sub-signal to a KR-level pass.
Anti-Patterns the Skill Detects
The skill scans for these and either flags or refuses:
- Retroactive target adjustment (we hit it because we changed the target) - document the change; grade against both definitions
- Retroactive scope shrinkage on a
committedorcompliance_or_safetyKR (committed to 3 healthcare audits, 1 audit completed, scored as "pass on in-scope") - refuse and mark not-yet-fully-observable - Average-the-guardrail-away (a failed guardrail dissolved into a high primary score) - separate the guardrail signal
- Aspirational-grading-of-committed (treating 0.7 as success on a committed KR) - refuse and explain
- Effort-equals-impact (initiative shipped, score did not move, score
Truncated for display — read the full file on GitHub.
Related Skills
last30days-skill
63.5kAI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
