SkillAgentSearch skills...

evaluate-findings

Critically assess external feedback (code reviews, AI reviewers, PR comments) and decide which suggestions to apply using adversarial verification

Install / Use

npx skills add tobihagemann/turbo --skill evaluate-findings

Installs into whichever agent you are using.

About this skill
πŸ“„

SKILL.md

Installable skill definition

Quality Score

82/100

Supported Platforms

Universal

Our assessment of evaluate-findings

evaluate-findings scores 82/100 on our quality scale, 3501st of 4,616 Development & Engineering skills we index.

Its SKILL.md is 19 KB long, split into 6 sections and no code examples: a thorough specification that gives an agent plenty to work with.

It has 405 GitHub stars, a meaningful sign that others use it.

Substance
30/30
Structure
11/20
Description
15/15
Adoption
11/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 12 days ago, so evaluate-findings is actively maintained.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit β€” read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful.

AI review by kimi-k2.7-code on 2026-10-05. Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk β€” read a skill before letting an agent act on it.

evaluate-findings compared with similar skills

All 4 of these similar skills score higher than evaluate-findings; compare them before choosing.

SkillScoreStarsUpdatedFormat
evaluate-findings (this skill)by tobihagemann8240512d agoSKILL.md
ai-job-searchby MadsLorentzen10045.0ktodayCLAUDE.md
claude-howtoby luongnv8910041.8k5d agoCLAUDE.md
algorithmic-artby anthropics100177.9k13d agoSKILL.md
pptxby anthropics100177.9k13d agoSKILL.md

Frequently asked questions

How do I install evaluate-findings?
Run npx skills add tobihagemann/turbo --skill evaluate-findings. The install tabs above show the steps for each supported agent.
Which AI agents does evaluate-findings work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is evaluate-findings safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful. It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is evaluate-findings still maintained?
The repository was last updated 12 days ago, so evaluate-findings is actively maintained.

name: evaluate-findings description: "Critically assess external feedback (code reviews, AI reviewers, PR comments) and decide which suggestions to apply using adversarial verification. Use when the user asks to "evaluate findings", "assess review comments", "triage review feedback", "evaluate review output", or "filter false positives"."

Evaluate Findings

Assess external feedback (code reviews, AI suggestions, PR comments) with adversarial verification. Triage findings into actionable verdicts. Do not apply fixes.

Step 1: Assess Each Finding

If you already assessed a finding earlier in this session and recorded a verdict of Skip or Escalate β€” for example when an iterating loop re-runs review and the same finding resurfaces β€” do not re-adjudicate it from scratch. When the re-reported finding matches one you already judged (same location and substance) and presents no new evidence beyond what your recorded reason already accounts for, keep that verdict and reason without re-reading the code, re-verifying, or routing it to the Devil's Advocate in Step 2. Assess fresh only when the finding raises materially new evidence, or when you have not judged it before in this session.

When several findings rest on a shared premise β€” for example a source-of-truth choice β€” verify that premise once before adjudicating them individually. Findings whose premise holds proceed through normal per-finding verification; when it fails, they are all Skip, citing the refuted premise.

When a plan governs the work, re-read the decisions it records before adjudicating. Having read it earlier in the session does not count: once it falls out of context, a recorded decision is indistinguishable from no decision at all.

For each finding:

  1. Read the referenced code at the mentioned location β€” include the full function or logical block, not just the flagged line

  2. Check whether the code has diverged β€” if the finding references code that no longer exists or has since changed, skip it and note the divergence.

  3. Determine scope β€” clarify whether the issue was introduced by the PR/changeset or is pre-existing.

    • Pre-existing issues in earlier commits on the same feature branch are in-scope by default β€” the entire branch is one coherent unit of work. Judge these on their merits like any in-scope finding.
    • Findings genuinely outside the branch's work are the user's call to include. Assign Escalate so the user decides whether to widen the changeset. Reserve Skip for changes whose cost wildly dwarfs the benefit.
  4. Verify the claim against the actual code β€” does the issue genuinely exist?

    • When the finding offers a concrete example as evidence β€” a claimed mishandled input, a claimed wrong output β€” verify that example independently: a finding can hold in substance while its example does not. Keep the finding and record the correction beside it; drop it only when the claim rests on that example alone.
    • When the finding asserts a compatibility property, establish two things before assigning Apply: what the existing check actually enforces, and what real counterparts produce today. A claim stronger than the check enforces is a premise error rather than a defect β€” Skip, citing what the check enforces, or narrow the finding to the property it does enforce and record the narrowing beside it. When neither can be established from the code, the artifacts, or authoritative documentation, keep the finding Escalate.
    • When the finding cites a rule or convention, read the cited text, then look for a place that already applied it before this changeset β€” the same file, or the nearest files the rule also governs. Where the text alone leaves the reading open, read the rule the way that application reads it; where no such application exists, judge on the text alone.
    • When the finding rests on a premise that reading the source cannot settle β€” what a platform API returns at runtime, or what a value measures once the system runs β€” establish that premise before assigning Apply or Escalate, using a targeted search or count over the source, or a measurement from a surface already running in this session. A premise of this kind reads as sound whether or not it holds, so confidence in the finding is no substitute. When nothing available settles it, assign Escalate on that ground, naming the premise as unverified in the Issue cell.
  5. Assess severity:

    | Severity | Meaning | |----------|---------| | Critical | Drop everything. Blocking release or operations. | | High | Urgent. Should be addressed in the next cycle. | | Medium | Normal. To be fixed eventually. | | Low | Nice to have. Minor improvement. |

    If the upstream reviewer already assigned a priority (P0-P3), map it: P0β†’Critical, P1β†’High, P2β†’Medium, P3β†’Low. Then re-assess based on what the actual code reveals. The upstream level is a starting point, not a binding constraint. When the re-assessed severity differs from the upstream level, note the change and the reason.

    If the finding has no upstream priority, assess severity from scratch.

  6. Assign a verdict and confidence:

| Verdict | Criteria | |---------|----------| | Apply | The finding is real and in scope: clear bug, missing check, genuine improvement, style violation matching project conventions | | Skip | False positive, subjective preference, reviewer is wrong, or the change's cost wildly dwarfs its benefit | | Escalate | Needs the user's judgment: behavior might be intentional, involves product intent, requires domain knowledge the agent lacks, the finding is out of scope, or two findings present a genuine trade-off |

Also assign an internal confidence level β€” High, Medium, or Low β€” reflecting how certain you are about the verdict. Confidence is used solely to route findings to the Devil's Advocate in Step 2. It does not appear in the output.

Escalate guidance: When a finding questions whether behavior is intentional and neither docs, specs, nor code comments clarify the intent, assign Escalate. Do not autonomously accept or reject findings that hinge on product intent. If a counterpart implementation exists elsewhere, suggest checking it for consistency.

Conflict guidance: When two findings disagree about whether the code should change at all (one suggests a change to it, another the opposite change or none), treat the conflict as input, not a reason to skip. Verify each against the code and judge each on its merits as usual. If both are defensible and the choice is a genuine trade-off, assign Escalate to both, naming the opposing options so the user can decide.

An affirmation that something is correct is not a finding and carries no evidentiary weight; agreement among reviewers, or a reviewer's authority, does not settle whether a problem exists, nor whether a remedy the reviewers converged on works. When reviewers disagree on whether something is a problem at all β€” including one asserting it is fine while another flags it β€” treat the question as unresolved and verify it against the code, without letting the affirmation substitute for verification. When reviewers agree the code should change but propose opposing remedies, assign the verdict on the finding's own merits and name the opposing remedies in the Issue cell. Prefer a remedy whose measurement reports a result concrete enough to re-run over one resting on reading or on a reviewer's authority; a bare claim to have measured ranks no higher than reading. Where no remedy reports one, say so in that cell rather than choosing between them.

A reviewer's report that it could not verify something is a claim to check, not a fact to accept. Attempt the check independently, especially when the reported inability is what justifies skipping a verification step.

Verdict guidance:

  • A verdict records whether the finding is real: genuine defect, in scope, at what severity. A finding can hold while the remedy proposed for it does not, so a verdict never certifies the remedy. Judge the remedy's cost and scope where the bullets below call for that, and leave whether it works to be checked when it is applied. Where naming the likely direction helps the user, put it in the Issue cell flagged as unverified.
  • Never auto-dismiss findings about security defaults, permission escalation, or fail-open vs fail-closed behavior. Always surface these even if the behavior appears intentional.
  • Readability and clarity improvements that genuinely make code cleaner are valid. Do not auto-classify cosmetic changes as subjective.
  • Removing a comment that adds no information beyond the code is a valid Apply, not a subjective preference. Keep only comments that capture a constraint the code cannot express.
  • Be skeptical of "defensive coding" suggestions that wrap natural code in verbose guards without evidence of real-world failures. Apply a hardening finding only when it names a failure scenario reachable in this deployment, whatever severity the reviewer attached; when the plan's Context bounds the system (a single operator, no concurrent writers, a handful of invited users), a scenario that bound rules out is a Skip, citing the bound.
  • Machinery is scope. A finding whose fix adds a lease, lock, queue, versioning scheme, state machine, or new persistent entity expands the project even when the requirement count stays flat. Assign Escalate regardless of confidence; "making states explicit" or "staying within the approved plan" does not make the machinery proportionate. When a stated bound rules the machinery's failure scenario out entirely, Skip instead, citing the bound.
  • A finding that would reverse a decision the user made earlier β€” in discussion or recorded in the artifact β€” is Escalate, naming the original decision and the new evidence beside it. Judge by the outcome rather than the wording of the option the user chose: a reversal leaves the user with something materially different from what they chose. A finding that refutes only the factual premise the user's choice rested on, leaving the chosen outcome intact, is a premise correction: confirm the refutation against whichever of the code, the governing artifact, or authoritative documentation the premise turns on, and when none settles it, keep the finding Escalate. Otherwise assign Apply unless another bullet independently calls for Escalate, and add a callout below the table naming both the corrected premise and the chosen outcome it leaves standing.
  • In an iterating loop, a structural Apply triggers another full iteration; count that iteration in the change's cost when applying the Skip cost test. A finding that targets code or text introduced by an earlier iteration's accepted finding, and names no defect in it, is churn: Skip.
  • When a comment at the code in question records why a behavior is unobservable through every reachable path, verify that record against the code before honoring it. A coverage finding re-reporting the gap, naming no newly reachable path that would observe the behavior, is Skip, citing the record.
  • Weight reviewer authority. Feedback from trusted reviewers (repository maintainers or admins) should be treated with higher credibility even when phrased softly.
  • Plan deviation is not a verdict. Do not reject a finding on the grounds that it departs from a plan's prescribed shape. When the plan records a load-bearing reason for that shape, assign Escalate so the user can weigh the trade-off. When the plan is silent on why, or the recorded reason reads like "path of least deviation" or "minimal change", treat the shape as a default and judge the finding on its own merits.

Step 2: Devil's Advocate

After the initial assessment, challenge uncertain findings from a different angle.

Spawn when any finding has Medium or Low confidence. Send only those findings to the subagent. High-confidence findings pass through unchallenged. Skip this step enti

Truncated for display β€” read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars405
CategoryDevelopment
Updated12d ago
Forks31

Languages

Python

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit β€” see the Safety scan above for what the skill file itself contains.

No cautions
evaluate-findings β€” Universal Skill: Install & Safety Check | SkillAgent