first-pass
Rules and checks that make Claude Code look around a change, not just at the lines it writes: ten questions before code, a reviewer that didn't write it, proof before done, bugs fixed as a class.
Install / Use
npx skills add joetawil7/first-passInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Our assessment of first-pass
first-pass scores 87/100 on our quality scale, 1227th of 2,864 Automation skills we index (top 43%).
Its SKILL.md is 24 KB long, well organised into 17 sections with 3 code examples: a thorough specification that gives an agent plenty to work with.
It has 147 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated today, so first-pass is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-01. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
first-pass compared with similar skills
All 4 of these similar skills score higher than first-pass; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| first-pass (this skill)by joetawil7 | 87 | 147 | today | SKILL.md |
| Agent-Reachby Panniantong | 100 | 87.2k | 16d ago | CLAUDE.md |
| rufloby ruvnet | 100 | 73.6k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.0k | 1d ago | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 9d ago | SKILL.md |
Frequently asked questions
- How do I install first-pass?
- Run
npx skills add joetawil7/first-pass. The install tabs above show the steps for each supported agent. - Which AI agents does first-pass work with?
- It is written for Claude Code, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is first-pass safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is first-pass still maintained?
- The repository was last updated today, so first-pass is actively maintained.
Skill content
View source on GitHubfirst-pass
first-pass makes a coding agent check the code around a change, not only the lines it writes, and prove its work before it calls anything done. It's a Claude Code plugin; Cursor, Codex and other agents get the same rules and skills (the hooks are Claude Code only). You set it up once, in the folder that holds your repos.
What it does
- Before code, the agent answers ten questions about the change against the real code:
what happens if it runs twice, stops halfway, or an outside service times out; which
other code uses the same data; what the user sees when it fails. Each answer is a
file:line, a test, or a gap you decide to accept. - While building, anything the task didn't ask for (a helper, a cap, a retry, an option) has to earn its place. The agent asks whether it needs to exist, leaves it out when nothing requires it, and lists it under "Not built" so you can ask for it. During the work, a review finding gets fixed only when it does real harm or stops the change doing its job; the rest come to you as one list at the end, to pick from.
- On the task you gave, the agent first writes the task's scope from your prompt and the
code around it: the feature, everything that uses the same data, the rest of the same
flow, and the steps that cancel or delete what it creates. Its searches and reviews stay
inside that scope. A serious problem it happens to see
elsewhere gets one line in the report (and one in that repo's list of known breaks when it
breaks one of its rules), and nothing else; anything the change itself breaks
counts as inside, wherever it is.
/first-pass:hulklifts the scope for one task, so it looks everywhere. - Before "done", a second agent that didn't write the code reviews it in a fresh context. "Done" needs a test that failed before the change and CI's checks passing in a clean checkout; anything less is reported as built, not done.
- On a bug, it fixes the whole class: it reproduces the bug, finds the same pattern in the task's scope, and adds the check that stops it coming back.
- On a pull request,
/first-pass:reviewchecks your own branch before you ask for review, or a teammate's PRs across several repos. Every finding comes with its proof or is marked unproven, and it never posts to GitHub./first-pass:review hulk <what>reviews with the scope lifted, so it looks everywhere. - In every session, hooks send back a "done" that has no evidence behind it (edits made through the shell count too), run each repo's own hooks from the shared folder, and say what drifted since setup.
- In replies (optional, you choose at setup): plain, short answers. The first line says what happened or what you need to do, in everyday words, without the names from inside the work, and the reply ends with only what you have to do or decide.
What it helps with
- Bugs next to the change: the other code path that writes the same field, the webhook that arrives twice, the error that shows up as an empty list.
- "Fixed" and "verified" that only meant "it compiles and the mocks pass".
- An agent reviewing its own work in the same context that wrote it.
- Many repos opened from one folder, where each repo's own rules and hooks don't load.
- Prompts like "be 100% sure", which change how sure the answer sounds, not what gets checked.
- Code nobody asked for: the extra cache layer, guard or option that becomes one more thing to maintain, and the next thing to break.
- Replies too long, or too full of the agent's own names for things, to follow once you've looked away.
/plugin marketplace add joetawil7/first-pass
/plugin install first-pass@first-pass
Then, in the folder where you start your sessions: /first-pass:setup-first-pass.
Why I made this
I built a product feature by feature with Claude Code. Every feature request ended with some version of "make sure the code is correct and bug free, cover all gaps, be 100% sure." Then I ran a full audit. It found 318 issues, 25 of them high severity.
When I sorted them, only 45 were mistakes in the lines being written. The rest were in the code around those lines:
| What went wrong | Issues | | ----------------------------------------------------------------- | -----: | | Another code path using the same data was not updated | 54 | | Runs twice, runs at the same time, or stops halfway | 54 | | An outside service fails or is slow, and the error is hidden | 45 | | A plain mistake in the code itself | 45 | | Time, units, rounding | 26 | | Scale: no limit, no index, lists that stop at one page | 25 | | Endings: cancel, delete, expire, reconnect, downgrade | 19 | | UI, help, legal or pricing text that says what the code doesn't do | 18 | | Pipeline: CI red, no tests against a real database | 17 | | Hostile user or uncapped cost | 15 |
(One private codebase, sorted by hand with one cause per issue. Your mix will differ.)
Several of these had been found by earlier audits and fixed. Each fix patched one spot, and nothing carried the lesson into the next session.
"Be 100% sure" changed how sure the answers sounded. It never made the agent open the other file that writes the same field, or ask what happens when the webhook arrives twice. So first-pass names those checks, and asks for proof before anything is called done:
- Ten questions before code (the pre-mortem): twice, halfway, outside call, failure is
not empty, neighbors, endings, money, hostile user, words, scale and time. Each answer is
a
file:line, a test, or "Not handled, because ___" for you to accept. - A reviewer that didn't write the code. Models are poor at catching their own mistakes
in the context that made them (Huang et al., 2024),
so the
breakeragent reviews each change in a fresh context, starting from the other code paths that touch the same data. - "Done" means a test that failed before the change, the breaker's review, and CI's own checks passing in a clean checkout. Short of that, it's reported as built, not done.
- Bugs get fixed as a class: reproduce, find the same pattern in the task's scope (the whole codebase when your own prompt asks for that fix, not when you pick it from a task's list), and add the helper, constraint or check that stops it coming back.
What's in it
| Piece | What it does | When |
| --- | --- | --- |
| premortem | The ten questions, answered against the code | Before code |
| breaker (agent) | Fresh-context review of the diff and of every other path touching the same data; concrete findings only | Before done |
| ship-check | The definition of done, ending in a report where every "Verified" line says what was run and its result | Before done |
| fix-the-class | Reproduce, name the class, search for it in the task's scope (everywhere when your own prompt asks for that fix, not when you pick it from a task's list), run the ten questions on the fix, fix or record each hit, make it hard to repeat | On any bug |
| setup-first-pass | Writes the rules once, a map of your repos, and a section per repo with its real commands, test limits and a drafted INVARIANTS.md | Once, then to update |
| habit-words | Reads what you typed in your recent sessions and maps words like "be 100% sure" to the checks they should mean | At setup, then when due |
| sharpen | Rewrites the prompt you type after it: numbered asks, habit words turned into checks, names and numbers kept exactly. Shows you the rewrite, then works from it | Only when you type it |
| review | Reviews your own branch before you ask for review (type nothing after it), or a teammate's PRs, several repos at once. Checks the change against its ticket, judges the failed checks and every Bugbot comment, runs the breaker, traces the other code that uses what changed, proves each finding or marks it unproven, and says what the merge needs and how to check the deploy. Reads GitHub and never posts; pushes a fix only on your yes | Only when you type it |
| jev | Sets up the optional Jev judge (TypeSafe's decision model), which ship-check asks whether a review finding is real harm, which small ones to fix now, and what proof a small fix needs | Only when you type it |
| hulk | Lifts the task's scope for one task: searches, reviews and bug hunts go across the whole codebase, and problems found anywhere are handled as usual (for a pull request: /first-pass:review hulk <what>) | Only when you type it |
| Hooks | Run each repo's own hooks from the main folder, send back a "done" with no evidence, say what drifted at session start, and hold sharpen's work until its rewrite is shown | Every session |
If you keep all your repos in one folder
A lot of us open one folder with every repo in it and start each session there. Claude Code
then finds agents, skills and hooks only in that folder and above it: a repo's own
.claude/agents, its .claude/settings.json hooks and its .cursor/rules/*.mdc imports
never load, and its CLAUDE.md loads only once a file in it is read. first-pass is built for
that:
- One set of rules at the root, in
AGENTS.md(imported byCLAUDE.md), loaded from the first message and updated in one place. - A section per repo, written from what setup finds in it: CI's exact commands, how to run one test, what the real tests need running, test limits, where its user-facing words live, and the same job done in two places. It loads when work reaches that repo.
- Each repo's own hooks still run. Setup lists them and you approve each one. An approval pins the hook and every file in its script's folder: if a pull changes one, the hook pauses until you approve it again, and you and the agent are both told.
- Drift is reported at session start: a repo's CI changed since its section was written, a new repo appeared, a hook was paused, a Cursor rule's copy fell behind, the rules are older than the plugin.
- One repo on its own works too: setup writes the rules and the section into that repo.
Your habit words
Most of us have words we type out of habit: "be 100% sure", "don't assume", "full review", "all fine, right?". They name no place to look, so the answer sounds more certain without anything more being checked.
habit-words reads what you typed in your last 20 Claude Code sessions, shows how often you
use each phrase, what went wrong after it, and what to say instead. Then it writes a short
block that maps each phrase to the checks it should trigger. You never have to type them
again, and if you do, they mean the checks, not a more confident tone.
What it reads and keeps:
- Only your own transcripts in
~/.claude/projects, and only what you typed, plus the end of the agent's reply before each prompt that pushes back on it. Tool output, pasted text, notifications and script-started runs are left out. - Anything that looks like a key, token, password, email or phone number is replaced before the text goes into a temp file (best effort, so the skill also never quotes one). The skill deletes the file when it's done.
- The block holds only your phrases and the checks, and never goes into a file your teammates share.
What it writes to your machine
AGENTS.mdandCLAUDE.mdat the root (the rules and, if you want them, the reply style and your habit words), and a section in each repo'sCLAUDE.mdor `AGEN
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
87.2kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
ruflo
73.6k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
Scrapling
85.0k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
