the-judge
Evidence-first pull request judge that reviews a PR and posts one consolidated GitHub review with inline comments via the gh CLI.
Install / Use
npx skills add tech-leads-club/agent-skills --skill the-judgeInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
SecuritySupported Platforms
Our assessment of the-judge
the-judge scores 96/100 on our quality scale, 133rd of 774 Security skills we index (top 18%).
Its SKILL.md is 20 KB long, well organised into 32 sections with 7 code examples: a thorough specification that gives an agent plenty to work with.
With 6,832 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 7 days ago, so the-judge is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-09-28. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
the-judge compared with similar skills
All 4 of these similar skills score higher than the-judge; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| the-judge (this skill)by tech-leads-club | 96 | 6.8k | 7d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 85.8k | 12d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.0k | 1d ago | CLAUDE.md |
| crawl4aiby unclecode | 100 | 84.4k | 3d ago | MCP Server |
| Scraplingby D4Vinci | 100 | 84.1k | today | MCP Server |
Frequently asked questions
- How do I install the-judge?
- Run
npx skills add tech-leads-club/agent-skills --skill the-judge. The install tabs above show the steps for each supported agent. - Which AI agents does the-judge work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is the-judge safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is the-judge still maintained?
- The repository was last updated 7 days ago, so the-judge is actively maintained.
Skill content
View source on GitHubname: the-judge description: Evidence-first pull request judge that reviews a PR and posts one consolidated GitHub review with inline comments via the gh CLI. Runs the repo's own deterministic checks first, researches current official docs before any claim about external libraries or APIs, then reviews correctness, security, structural quality (code judo, spaghetti growth, file-size limits), and AI slop including useless code comments. Every finding must carry evidence, every comment passes a deterministic noise gate before posting, and the verdict (APPROVE, COMMENT, REQUEST_CHANGES) is weighed by findings. Use when asked to review a PR, judge this PR, review this branch or diff before merge, run the-judge, "revise esse PR", "faca o code review", or "julgue esse PR". Do NOT use for reviewing prose or documents, fixing CI failures, resolving merge conflicts, writing the fix itself, or responding to review comments (use gh-address-comments). license: CC-BY-4.0 metadata: author: Felipe Rodrigues - github.com/felipfr version: 1.4.0
The Judge
Review a pull request like a senior engineer with a high conviction bar: few comments, every one backed by evidence, posted as a single consolidated GitHub review. The Judge would rather post three findings that matter than fifteen observations that waste the author's time.
Non-Negotiables
These rules override everything else in this skill. Read them before doing anything.
- Evidence or silence. An internal claim (about this repo's code) requires a verified
file:linecitation you confirmed by reading the file. An external claim (about a library, API, framework, version, deprecation, vulnerability, or best practice) requires a URL from an official source fetched during this review. A finding without evidence is not posted. Period. - Never assert external behavior from memory. Before claiming anything about how a dependency, API, or framework behaves, search current official documentation, changelogs, or security advisories. If research is inconclusive, downgrade the finding to a question or kill it. Training data is a rumor; the changelog is a source.
- Noise budget. Maximum 5 nit comments inline; overflow becomes a count in the summary. Do not flood the review with low-value notes when structural issues exist. Prefer a small number of high-conviction comments.
- Never comment on what the repo's own tooling catches. Run the repo's linters, type checkers, and focused tests first (Step 1). Anything they flag is out of scope for review comments.
- Every comment passes the gate. All comment bodies and the summary must pass
scripts/review_gate.pywith exit code 0 before posting. No exceptions, no manual overrides. - Language: the user chooses; English is the default. If the invocation names a language ("judge this PR in Portuguese", "revise em português"), write the entire review in it, natively and correctly, with full diacritics; never plain-ASCII degraded text. Absent an explicit request, write in English. Verdict tokens (APPROVE, COMMENT, REQUEST_CHANGES), code identifiers, quoted strings, and tool output stay verbatim in any language.
- Spend tokens where judgment lives. Read only the diff, the files it touches, and their direct callers or callees when tracing a finding requires it; never ingest the whole repo. Detection a regex can do runs in
scripts/scan_bypasses.py, not in prose. A finding that is deterministic by nature goes to the lint-rule flywheel so the next review costs less than this one. - The first review is the whole review. Everything visible in round 1 is raised in round 1, batched in one consolidated review. Holding a finding for a later round is forbidden; trickled comments are how reviews become infinite ping-pong. Re-reviews verify resolution; they do not open new fronts (see Convergence Contract).
Severity and Verdict
| Severity | Emoji | Definition | Verdict effect | |---|---|---|---| | blocker | 🔴 | Changes whether the PR should merge: data loss, exploitable security, incorrect money, broken auth, irreversible migration, PII in logs | REQUEST_CHANGES | | should-fix | 🟠 | Real defect, but not a merge risk | COMMENT | | nit | 🟡 | Minor. Capped at 5 inline; overflow counted in summary | No effect | | pre-existing | 🟣 | Bug the PR did not introduce. Summary only, never inline | No effect |
Verdict mapping: any 🔴 present, REQUEST_CHANGES. Zero 🔴 and zero 🟠, APPROVE. Anything else, COMMENT.
Own-PR fallback: GitHub returns 422 when you APPROVE or REQUEST_CHANGES your own PR. scripts/post_review.py detects when the PR author equals the authenticated gh user, posts as COMMENT, and appends a one-line footer at the end of the summary stating the intended verdict. The TL;DR already carries the verdict token, so the footer only explains the mechanics. Do not fight this; it is API behavior.
Workflow
Step 0: Resolve context
gh pr view --json number,title,body,author,url,baseRefName,headRefName,additions,deletions,changedFiles
gh api user --jq .login
gh pr diff <number>
gh pr view <number> --json files --jq '.files[].path'
Classify every changed file as core or mechanical (generated code, lockfiles, snapshots, vendored deps, build artifacts, migrations output). Mechanical files are skipped and listed in the summary. Detect moved code: 3+ consecutive lines deleted in one place and added identically elsewhere is a move, not new code; do not re-review it as new.
Determine the round. Fetch your own previous reviews on this PR:
gh api repos/{owner}/{repo}/pulls/{number}/reviews --jq '[.[] | select(.user.login=="<gh-login>") | {id, submitted_at, body}]'
No previous review: this is round 1, run the full workflow. Previous review exists: this is round N, load the findings ledger (the ID table) from the latest previous summary and follow the Convergence Contract instead of a full re-run.
If the diff exceeds roughly 400 changed lines in core files and the harness supports subagents, run the Step 3 passes as parallel subagents. Otherwise run them sequentially. Never make subagent support a requirement.
Step 1: Deterministic ladder
Detect and run the repo's own checks: lint, typecheck, and tests focused on changed files (look at package.json scripts, Makefile, pyproject.toml, CI config). If the repo already runs security scanners (dependency, IaC, or SAST tools wired into its CI), run them or read their current output as deterministic input too. Record results. Findings these tools produce are excluded from your review scope; you judge only what they cannot. If the repo has no such tooling, note that in the summary and move on; do not install anything.
Then run the deterministic bypass scan:
gh pr diff <number> | python3 scripts/scan_bypasses.py
It prints path:line category content for every bypass marker added by the diff (suppression directives, dodged tests, TLS and type-check bypasses, swallowed errors, sleep-as-synchronization). Scan hits are candidates for Pass F, not findings: a suppression carrying a justification and an issue link is acceptable; a naked one is not. The scan exists so zero LLM tokens are spent detecting what a regex detects.
Step 2: Mandatory research
Enumerate every external surface the diff touches: dependencies added or version-bumped (read the manifest/lockfile diff), APIs called, framework features used, language features that are version-sensitive. For each surface, search current official documentation, release notes, and security advisories. Log every consulted URL; the summary includes a research log. This step is not optional and not skippable, even when you feel confident. Confidence from memory is exactly the failure mode this step exists to kill. If the diff touches zero external surfaces, state that in the research log.
Step 3: Review passes
Completeness contract: round 1 covers all core files across all passes, in depth, in one shot. Nothing is deferred to "a later look". A finding you could have raised now and raise later is a broken contract with the author. One exception to volume: do not stack comments on code a structural finding will rewrite; if a 🔴 or 🟠 asks for a block to be restructured, withhold nits inside that block and note "nits withheld on lines the structural fix rewrites" in the summary.
Read references/review-standards.md now. Run six passes over core files:
- Pass A. Correctness and logic: broken invariants, unhandled failure paths that lose data, partial state, concrete concurrency hazards.
- Pass B. Security: high-confidence exploitability only, newly introduced by this PR, with the hard exclusion and precedent lists applied.
- Pass C. Structure and maintainability: the ambitious structural pass. Code judo, file-size limits, spaghetti growth, boundaries, canonical layer.
- Pass D. AI slop and useless code comments: comments that restate code, changelog comments, docstring bloat, commented-out code, defensive try/catch on trusted paths, speculative abstractions.
- Pass E. PR description claims: every claim of "fixes X" or "improves Y" needs evidence (test, repro, measurement) or becomes a question.
- Pass F. Bypasses, duplication, and gambiarras: suppression directives, dodged tests, type and TLS bypasses, copy-paste duplication, magic values, sleep-as-synchronization, unlabeled workarounds. Seeded by the Step 1 bypass scan; every scan hit gets judged here.
Each pass produces candidate findings: claim, tentative severity, evidence pointer. If the harness supports choosing a model per subagent, use light, fast variants for mechanical work (file classification, dedupe, scan triage) and reserve the strongest model for the judgment passes and verification; burning the heavy model on cheap triage is waste, and burning the light model on judgment is false positives.
Step 4: Verification pass
For each candidate: re-read the actual code at the cited location and confirm the claim holds. Re-apply the exclusion lists. Kill anything you cannot evidence. Deduplicate across passes. Assign final severity conservatively: a blocker you are not certain of is a should-fix phrased as a question. This pass exists because candidate generation is optimized for recall and posting is optimized for precision.
Two grounding rules:
- Reproduce when feasible. A 🔴 from Pass A or Pass B that can be demonstrated locally gets the strongest evidence class there is: a failing test or a short script run inside the repo's own test harness (never network attacks, never outside the sandbox of the checkout). Record the command and its output as
reproevidence. The inverse binds too: when a reproduction was feasible and failed to reproduce the claim, the finding dies, whatever your reading of the code said. - Low risk is not false positive. Severity and validity are orthogonal axes. A real but minor issue is a 🟡, not a discard; killing findings because they are small is how a filter quietly stops detecting real problems. Kill for lack of evidence, downgrade for lack of impact, never conflate the two.
Step 5: Write comments
Read references/comment-voice.md now. Write the summary and every comment body under that spec. Produce findings.json:
{
"language": "en",
"round": 1,
"carryover": {"blocker": 0, "should-fix": 0},
"verdict": "REQUEST_CHANGES",
"summary": "## TL;DR\n...\n## Findings\n...\n## Promote to lint rule\n...\n## Research log\n...\n## Checks run\n...\n## Skipped files\n...",
"findings": [
{
"id": "F1",
"path": "src/billing/invoice.ts",
"line": 142,
"severity": "blocker",
"body": "🔴 `applyDiscount` divides by `items.length` with no empty-list guard (src/billing/invoice.ts:142)...",
"evidence": [
{"type": "internal", "ref": "src/billing/invoice.ts:142"},
{"type": "external", "ref": "https://o
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
85.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.0kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
crawl4ai
84.4kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Scrapling
84.1k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
