SkillAgentSearch skills...

research-implement-feature

Build a working artifact from a plain \"implement X for me\" request: a running end-to-end spine first, then one feature per rung, with every under-determined decision written to an assumption ledger BEFORE the code that depends on it and a cross-model sweep for the ones that slipped through undecla…

Install / Use

npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill research-implement-feature

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

98/100

Category

Design

Supported Platforms

Universal

Tags

Our assessment of research-implement-feature

research-implement-feature scores 98/100 on our quality scale, 8th of 196 Design skills we index (top 5%).

Its SKILL.md is 31 KB long, well organised into 29 sections with 7 code examples: a thorough specification that gives an agent plenty to work with.

With 16,644 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
18/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 9 days ago, so research-implement-feature is actively maintained.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

research-implement-feature compared with similar skills

All 4 of these similar skills score higher than research-implement-feature; compare them before choosing.

SkillScoreStarsUpdatedFormat
research-implement-feature (this skill)by wanshuiyin9816.6k9d agoSKILL.md
algorithmic-artby anthropics100177.9k5d agoSKILL.md
pptxby anthropics100177.9k5d agoSKILL.md
designby nextlevelbuilder100130.2k6d agoSKILL.md
ui-ux-pro-maxby nextlevelbuilder100130.2k6d agoSKILL.md

Frequently asked questions

How do I install research-implement-feature?
Run npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill research-implement-feature. The install tabs above show the steps for each supported agent.
Which AI agents does research-implement-feature work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is research-implement-feature safe to use?
It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is research-implement-feature still maintained?
The repository was last updated 9 days ago, so research-implement-feature is actively maintained.

name: research-implement-feature description: "Build a working artifact from a plain "implement X for me" request: a running end-to-end spine first, then one feature per rung, with every under-determined decision written to an assumption ledger BEFORE the code that depends on it and a cross-model sweep for the ones that slipped through undeclared. Use when user says "给我实现", "implement X", "帮我做一个能跑的", "先搭个原型再加功能", "build this feature", "prototype then extend", or hands over a capability description rather than an experiment plan." argument-hint: "[what-to-build] [— effort: lite|balanced|max|beast] [— ask: never|semantic] [— base repo: <url>]" allowed-tools: Bash(*), Read, Write, Edit, Grep, Glob, Skill, AskUserQuestion, mcp__codex__codex, mcp__codex__codex-reply

Research Implement: Feature

Build: $ARGUMENTS

This skill exists for one request shape — "just implement X for me" — where the user has a capability in mind, not an experiment plan, and does not want to be interviewed about it first.

It resolves that request the only honest way: stay autonomous, stop being silent. The skill never blocks to ask permission; it declares every decision the request left open, in a ledger, at the moment it makes it, and then a different model family goes looking for the ones it forgot to declare.

Two invariants

  1. Declare before you act. The instant a decision is under-determined by the request and changes an interface or a meaning, it gets a ledger row — before the code that depends on it exists. A ledger reconstructed at the end of the run is not a ledger, it is a changelog, and it systematically omits exactly the assumptions the author stopped noticing.

    Under ASK=semantic, this invariant strengthens to ask before you act for the semantic class: the ledger row is the unit of ambiguity, so a row that would have been written silently is a question that gets asked first.

  2. Spine before features. Rung F0 is a walking skeleton: the thinnest path from real entry point to real artifact, with stubs inside. It must run before any feature is added. Features are then added one rung at a time, each with its own acceptance check, each leaving every earlier rung green.

Scope boundary

| The ask | Route | |---|---| | "implement X" / "build me something that does X" / "prototype then extend" | this skill | | "find me a research direction and take it to a paper" | /research-pipeline | | "I have EXPERIMENT_PLAN.md — run the campaign, deploy to GPU" | /experiment-bridge | | "sweep these parameters / find the best config" | /dse-loop | | "launch what is already written" | /run-experiment | | "do these results support the claim?" | /result-to-claim |

Relationship to /research-pipeline

/research-pipeline answers "what should we research?" and decides the question for you. This skill answers "build the thing I already decided on" and decides nothing of consequence without writing it down. Different input contracts, so they are different entry points rather than a mode flag — but they compose: a pipeline run may delegate its build stage here instead of inlining implementation, and inherits the ledger as a result.

If the target decomposes into more than the rung budget below, the scope is too large for one run. Cut to the MUST rungs and record the rest under Deferred in the build note — do not quietly grow this skill into a system build.

Constants

  • EFFORT = balanced — Work intensity per shared-references/effort-contract.md. Override: — effort: max.

    | | lite | balanced | max | beast | |---|---|---|---|---| | Rung budget (Phase 1) | 3 | 5 | 8 | 12 | | Fix attempts per rung (Phase 3) | 3 | 5 | 8 | 12 | | Silent-assumption sweep rounds (Phase 4) | 1 | 2 | 2 | 3 | | Reuse survey depth (Phase 0) | local grep | local + ecosystem | + reference impl | + fetch & diff reference impl |

    EFFORT never lowers the reviewer tier — a hard invariant of the effort contract.

  • ASK = never — Interaction mode: which ambiguities are put to the author before they are acted on.

    | — ask: | Asks about | Blocking? | For | |---|---|---|---| | never (default) | nothing — declare and proceed | no | unattended runs, overnight, /loop, a request you want executed not discussed | | semantic | semantic rows only | at batch points | you trust the small calls, you want a say in what the results will mean |

    ASK never changes what lands in the ledger — only who decided each row. Every row records its Source, so the record is complete in both modes.

  • ASSURANCE — derived from EFFORT per the effort contract (lite/balanced → draft, max/beast → submission). Governs whether Phase 4 blocks. Override: — assurance: submission.

  • BASE_REPO = false — Repo URL to build on top of. When set, clone first and implement inside it, matching its conventions. When false, extend the current project or create files in it.

  • Output language — follow shared-references/output-language.md. Code, paths, config keys and ledger IDs stay English regardless.

Interaction rule (HARD CONSTRAINT)

Resolve ASK once from $ARGUMENTS before Phase 0 and hold it for the run.

ASK=never — non-blocking

Runs end-to-end with zero external approval: no AskUserQuestion, no "should I…", no "please confirm", no waiting. Framework choice, file layout, whether to overwrite, whether to install a dependency, which default to pick — all decided here, and the consequential ones logged. The author reviews the ledger and the diff after the run.

Autonomy is not permission to be vague. Every decision you make instead of asking that changes an interface or a meaning is a decision you owe the author a row for.

ASK=semantic — blocking at batch points

The run stops and ends the turn at a batch point and resumes only on an explicit reply. Never implement this as "ask, then continue if no answer arrives" — once the turn ends, silence cannot resume the run.

Batch points (the only places questions are allowed): B0, end of Phase 0, before the ladder is built · B1..Bn, start of each rung, before that rung's code · Bd, a debugging fork where the fix itself is a semantic choice ("shapes don't match: pad left or right?").

Collect the batch and ask it in one call, never one question at a time. The chosen default is always option 1, labelled (default), so accepting everything as-is is one keystroke and produces exactly what ask: never would have. An answer of "you decide" (or an Other reply that declines to choose) falls back to that default, records Source: default (deferred_to_author), and is never re-asked. A batch point with nothing in it is skipped silently — it is not a checkpoint to announce.

Do not combine ask: semantic with /loop, CronCreate, or any overnight cadence. A blocking gate on an unattended run is a run that did nothing. Detect this at Phase 0 — if there is no interactive author, say so and stop rather than silently downgrading to never.

Acceptance-gate provenance

Per shared-references/acceptance-gate.md:

| Gate | Type | Who signs off | |---|---|---| | "the F0 spine ran end-to-end" | A | shell exit code + test -f on the artifact | | "rung Fi's acceptance check passed" | A | that rung's one-command check, exit code | | "no earlier rung regressed" | A | the accumulated check suite, exit code | | "fix budget / sweep-round budget exhausted" | A | a counter | | "the code silently assumes something the ledger does not declare" | B | Codex (Phase 4) — a different model family reads the diff cold | | "the implementation is correct / the method works" | B | out of scope here — belongs to /experiment-audit and /result-to-claim |

The terminating condition of the build loop is Type-A only. On a green run this skill says "the spine runs and every MUST rung's check passed". It never says the implementation is correct, the method works, or the numbers mean anything — a passing smoke test is an execution fact, not a result.

The one Type-B gate it does own is Phase 4, and it is owned for a reason: "what did I assume without saying so" is precisely the question an author cannot answer about their own work, because the assumptions they absorbed are the ones they stopped seeing. That needs a reader from a different family, not a second pass by the same one.

Artifacts

All under implement-stage/ (stage-scoped per shared-references/output-manifest.md; stage = implementation):

| File | Written | Contents | |---|---|---| | SPEC.md | Phase 0 | the request, restated as target / inputs / outputs / success command / base commit / scope cuts | | ASSUMPTIONS.md | Phase 0 onward, continuously | the ledger — one row per under-determined decision that changes an interface or a meaning | | BUILD_NOTE.md | Phase 1 onward | the ladder, the per-rung run record, deferred rungs, and blockers — one file | | SILENT_ASSUMPTION_SWEEP.json | Phase 4 | the cross-model verdict — the inspectable receipt that the acquittal was external |

Create implement-stage/ if absent. Do not create a MANIFEST.md — this run produces well under the 15-artifact threshold.

The assumption ledger

Schema

implement-stage/ASSUMPTIONS.md:

# Assumption Ledger — <target>
<!-- ASK mode: never | semantic -->

| ID | Under-determined by the request | Chosen | Class | Source |
|----|--------------------------------|--------|-------|--------|
| A-001 | request says "on the benchmark", does not say which split | validation | semantic | user |
| A-002 | no tokenizer named | reuse the repo's existing `BPE-32k` | interface | default |

## Notes

Prose, only where a decision is genuinely contested: the alternative that was
rejected and why, what reversing it would cost, and the one-line override.

- **A-001** — `test` is the held-out split and `train` leaks; `validation` is the
  only choice that leaves the number meaning what a reader assumes. Reversing it
  is one line in `configs/eval.yaml`.

Which decisions get a row. Only interface and semantic ones:

| Class | Means | Handling | |---|---|---| | interface | changes call sites, configs, or artifact schemas | ledger row + named in the final report | | semantic | changes what a result would MEAN — metric definition, eval split, normalization, what counts as a baseline, what the null hypothesis is | ledger row + its own block at the top of the final report + never summarized away + the only class ask: semantic gates on |

Naming, log format, file layout, and anything internal to one module that is invisible at its interface: just make the call. They do not get rows. A ledger that logs variable names buries the two rows that actually decide what the work will later claim, and turns every decision into a form.

The semantic class is the whole point. An undeclared interface assumption costs a refactor. An undeclared semantic assumption is how an implementation quietly decides what the research will later claim.

Source records who decided the row:

| Source | Means | |---|---| | user | the author was asked at a batch point and chose this | | default | this skill chose it — ASK did not cover the class, or the row was written after the batch point had passed | | default (deferred_to_author) | the author was asked and answered "you decide" | | sweep | Phase 4 found it undeclared and it was added retroactively |

Under ask: semantic, a plain default row in the semantic class is exactly an ambiguity the skill did not recognise as an ambiguity in time to ask about it — which is the most interesting row

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars16.6k
CategoryDesign
Updated9d ago
Forks1.4k

Languages

Python

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions