SkillAgentSearch skills...

gm-evaluate

Evaluate the current tag's quality: enforce the playable-closed-loop gate, maintain a single cross-tag e2e/ suite that always reflects the current game (add tests for new mechanics, prune tests for mechanics this tag deliberately removed), and reason about gameplay quality.

Install / Use

npx skills add RandallLiuXin/GodotMaker --skill gm-evaluate

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

88/100

Category

Other

Supported Platforms

Universal

Tags

Our assessment of gm-evaluate

gm-evaluate scores 88/100 on our quality scale, 83rd of 241 Other skills we index (top 35%).

Its SKILL.md is 18 KB long, well organised into 12 sections with 1 code example: a thorough specification that gives an agent plenty to work with.

It has 549 GitHub stars, a meaningful sign that others use it.

Substance
30/30
Structure
17/20
Description
15/15
Adoption
12/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 16 days ago, so gm-evaluate is actively maintained.
  • No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
  • Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

gm-evaluate compared with similar skills

All 4 of these similar skills score higher than gm-evaluate; compare them before choosing.

SkillScoreStarsUpdatedFormat
gm-evaluate (this skill)by RandallLiuXin8854916d agoSKILL.md
algorithmic-artby anthropics100177.9k11d agoSKILL.md
pptxby anthropics100177.9k11d agoSKILL.md
designby nextlevelbuilder100130.2k12d agoSKILL.md
ui-ux-pro-maxby nextlevelbuilder100130.2k12d agoSKILL.md

Frequently asked questions

How do I install gm-evaluate?
Run npx skills add RandallLiuXin/GodotMaker --skill gm-evaluate. The install tabs above show the steps for each supported agent.
Which AI agents does gm-evaluate work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is gm-evaluate safe to use?
It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is gm-evaluate still maintained?
The repository was last updated 16 days ago, so gm-evaluate is actively maintained.

name: gm-evaluate description: | Evaluate the current tag's quality: enforce the playable-closed-loop gate, maintain a single cross-tag e2e/ suite that always reflects the current game (add tests for new mechanics, prune tests for mechanics this tag deliberately removed), and reason about gameplay quality. Independent from the build process — fresh perspective on the final product. Explicit invocation only — use /gm-evaluate. disable-model-invocation: true

GodotMaker Evaluate

$ARGUMENTS

You are an independent game quality evaluator. You have NOT seen the build process. You only care about the final result for the current tag: does the game (as it stands at this tag) deliver the current Playable Unit and every mechanic the project has shipped so far — including the ones this tag adds, and the inherited ones from previous tags that should still work?

E2E tests live in a single e2e/ directory that always reflects the current state of the game. There is no per-tag e2e partitioning: when a tag adds a mechanic you add a test; when a tag deliberately removes a mechanic the corresponding refactor task in PLAN's Main Build prunes the test in the same change. You maintain e2e/ so it matches the union of every still-supported mechanic listed across the current PLAN's Tag Mechanics + Inherited Mechanics.

Session Setup

FIRST ACTION — before anything else: Write evaluate to .godotmaker/current_role.

Permission: You can write to e2e/, .godotmaker/evaluation.json, and append to .godotmaker/stage.jsonl (plus .godotmaker/current_role set during Session Setup). All other files are read-only.

Resume Check

Read .godotmaker/stage.jsonl (treat as empty if missing) — each line is {"role": X, "ts": Y}.

  • If no event with role == "verify" exists anywhere in the file → STOP. Tell user to run /gm-verify first.
  • If PLAN.md is missing the **Tag:** header → STOP. Tell user the file is stale and to re-run /gm-gdd to regenerate it for the current tag.
  • If the last event has role == "evaluate" AND .godotmaker/evaluation.json exists → STOP. Tell the user:

    "Evaluate already ran at {timestamp} with no verify since. Recommended next: /gm-accept (if approved) or /gm-fixgap (if rejected). If you need to redo this step or have other plans, just tell me."

  • If the last event is role == "fixgap" with exactly outcome == "handoff", next_role == "evaluate", and reason == "evaluator_owned_e2e" → read the handoff notes and referenced runtime evidence; repair the e2e/ scenario/assertion/capture timing; re-run affected checks; write a new evaluation.
  • Any other fixgap event carrying outcome → STOP. Ask the user to run /gm-fixgap.
  • Otherwise → proceed (evaluate is naturally re-invoked after each verify pass).

Resolve godot binary

Read godot_path from .claude/godotmaker.yaml and substitute it verbatim for <godot_path> in every godot --headless … command below. The path was validated at publish time and is the source of truth for which Godot binary this project uses.

If .claude/godotmaker.yaml is missing the godot_path field, fall back to plain godot (PATH lookup). If THAT also fails, STOP and tell the user Godot binary not configured — re-run tools/publish.py to set godot_path in .claude/godotmaker.yaml. Do NOT spelunk through PATH directories or guess install locations.

Evaluation Process

Phase 1 — Understand Requirements

Read in order:

  1. PLAN.md — extract Tag: header (call it <Tag>), Tag Mechanics list, Inherited Mechanics list, Playable Unit, Main Build refactor tasks (the latter tells you which prior-tag mechanics this tag intentionally removes)
  2. GDD.md — design intent (north star); cross-reference Tag Mechanics against the relevant GDD sections
  3. STRUCTURE.md — current tag's ECS architecture
  4. SCENES.md — current tag's scenes
  5. ASSETS.md — cross-tag asset manifest
  6. ROADMAP.md — confirm <Tag> is the entry being worked on (it should be the earliest entry without a git tag)

Build a single expected-mechanics checklist = (every [<Tag>-MN] from Tag Mechanics) ∪ (every [<prev>-MN] from Inherited Mechanics). This is the union of mechanics the game must currently support. The corresponding test files in e2e/ must cover this checklist exactly — no more, no less.

Build a playable-unit checklist from PLAN.md Playable Unit: player experience, unit outcome, scenes involved, and every row in the per-mechanic playability table.

Key playable_unit.rows by mechanic id, for example v0.1.0-M1.

Phase 2 — Maintain the e2e/ suite

E2E tests live in a flat e2e/ directory (no per-tag subdirectories). Each test file is named after the mechanic id it covers, e.g. e2e/test_v0.1.0_M1_wasd_movement.gd — the mechanic id in the filename keeps the test→ID mapping mechanical and stable as later tags inherit it.

  1. Read .claude/skills/godot-e2e/SKILL.md for the API.
  2. Confirm e2e/conftest.py exists at the e2e root (created by gm-scaffold).
  3. Add tests for new Tag Mechanics: for each [<Tag>-MN] in PLAN.md that does not yet have a test file in e2e/, write e2e/test_<tag_slug>_M<N>_<mechanic_slug>.gd (or .py). The test must assert the observable behaviour named in the mechanic line, not internal state.
  4. Add or update Playable Unit coverage tests: write e2e/test_<tag_slug>_playable_unit_<slug>.gd (or .py) files until every Playable Unit table row is covered. Each covered row must exercise player-facing runtime behavior, assert the expected effect, and capture or reference the required visible content. If the row names a completion/fail/exit state, the test must reach it through play.
  5. Verify Inherited Mechanic tests still exist: for each [<prev>-MN] in PLAN.md's Inherited Mechanics, the corresponding test file from when that prior tag shipped must still be in e2e/. If a file is missing (e.g. accidentally deleted), restore it by reading docs/tags/<prev>/PLAN.md and re-implementing the test.
  6. Prune tests for removed mechanics: if PLAN's Main Build has a refactor task that removes a prior-tag mechanic (and that mechanic id therefore does NOT appear in this tag's Inherited Mechanics list), delete the corresponding e2e/test_*.gd file. Removal is intentional, refactor task is the audit trail.
  7. Add scene-transition tests for new scenes added in this tag.
  8. Run the full suite: godot-e2e e2e/ -v
  9. Fix test bugs (wrong node paths, timing issues) — but do NOT fix game bugs; those are Phase 3+ findings.

E2E repair boundary:

  • E2E tests verify observable gameplay requirements. Do NOT prescribe fixgap's implementation strategy.
  • When a state cannot be reached reliably in E2E, record the observed gap and request the deterministic test interface needed to exercise it.
  • Test interfaces include simulate_* methods, scene setup helpers, fixed seeds, public state setup, or debug-safe setup paths that call real runtime code.
  • Do NOT ask fixgap to change normal gameplay behavior, balance, progression, content, or timing only to satisfy a test assertion.
  • If normal gameplay itself violates GDD/PLAN, cite the design source and record the gameplay failure.

When requesting a test interface, write both evidence entries:

  • observed_gap: <observable state or behavior not proven by E2E>
  • requested_test_interface: <bounded setup or simulate interface needed>

After this phase the e2e/ directory must contain exactly one test file per mechanic id in the expected-mechanics checklist (Phase 1), Playable Unit coverage for every Playable Unit table row, plus scene-transition tests. Stale files for mechanics that no longer appear anywhere are a Phase 3 critical_issue.

Before completing /gm-evaluate, write one playable_unit.rows entry for every PLAN Playable Unit row. Each entry must include result, test, and non-empty evidence; the referenced test file must exist. For approve, every row must be pass.

Phase 3 — Mandatory Checks

All of these must pass for result == "approve". Failure of any is a critical_issue.

Playable closed loop (composite hard gate):

  1. Builds clean: "<godot_path>" --headless --quit 2>&1 — zero ERROR lines.
  2. Boots into main scene: project.godot points to the right entry scene; the entry scene loads without crash (confirm via E2E).
  3. Playable Unit coverage passes: every Playable Unit table row has passing E2E coverage for the player operation/content, expected effect, and required visible content.
  4. Completion/fail/exit is reached through play: every completion/fail/exit state named in the Playable Unit is triggered by E2E through normal play. Static code evidence is not enough.

Mechanics gate (covers both new and inherited): 5. Every entry in the expected-mechanics checklist has a corresponding test in e2e/ AND that test passes. Each PASS/FAIL recorded against the mechanic id. A failing inherited test is just as critical as a failing tag test — both block approval. 6. The e2e/ directory must NOT contain test files for mechanic ids absent from the checklist (orphan tests). Fix by either re-adding the missing mechanic to PLAN's Inherited Mechanics, or pruning the orphan test (whichever matches actual game state).

Visual cross-check (per scene listed in SCENES.md): 7. Capture screenshots under e2e/screenshots/. Use game.screenshot("e2e/screenshots/scene_{name}.png") for static scenes. For scenes with motion/animation, capture a frame sequence per .claude/skills/screenshot/SKILL.md § "Frame Sequence for VQA Dynamic Mode". Treat e2e/screenshots/ as latest-run output only. 8. Inspect the captured screenshots by dispatching a subagent to run the visual-qa skill in Question mode. Do not compare screenshots against references/scene_{name}.png.

Visual binding preflight. Before calling visual-qa, check the scene's Asset bindings rows. Each non-procedural / non-UI text / non-not required this tag binding must have:

  • a concrete asset_name / path value;
  • a matching ASSETS.md Asset Table row;
  • a matching ASSETS.md Visual Asset Contract row;
  • a non-empty Runtime Size;
  • an existing file when the row status means the asset should be on disk.

If the scene has no Asset bindings section or ASSETS.md has no Visual Asset Contract section, record missing visual contract for <scene> in major_issues, then continue VQA with the scene Acceptance criteria and mechanic fallback context. If the sections exist but a current-tag binding is incomplete, record a critical_issue, set this scene's visual_checks.<scene>.result to "fail", note the exact missing binding in visual_checks.<scene>.notes, and skip visual-qa for that scene.

For not required this tag, require a deferral reason in the Visual Contract or Readability Requirement text. Missing deferral reasons are incomplete bindings.

Question construction. Pull the Acceptance criteria block from SCENES.md for this scene. Add the scene's Asset bindings rows and matching ASSETS.md Visual Asset Contract rows. If the block is absent, fall back to the mechanic ids from PLAN.md Tag Mechanics + Inherited Mechanics that this scene exercises, each with its one-line description. Ask whether the screenshot or frame sequence satisfies the required visible content, readability, layout, and motion/animation requirements. For deterministic setup screenshots, add Visible state only; do not infer prior play history..

Atlas misuse check. When a scene's Asset bindings reference a region_atlas (or a single element sourced from a grid_sheet), add to the question whether each element shows only its intended region/sprite and not the whole atlas or sheet. A visible full atlas or sheet where a single button, icon, prop, or FX s

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars549
CategoryOther
Updated16d ago
Forks50

Languages

Python

Trust signals

88/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 medium