SkillAgentSearch skills...

post-patch-validation

Validates security patches with reproducible baseline-versus-patched evidence, including original exploits, root-cause variants, behavior preservation, regressions, and newly introduced security failures.

Install / Use

npx skills add trailofbits/skills --skill post-patch-validation

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

95/100

Category

Security

Supported Platforms

Universal

Our assessment of post-patch-validation

post-patch-validation scores 95/100 on our quality scale, 232nd of 774 Security skills we index (top 30%).

Its SKILL.md is 15 KB long, well organised into 11 sections with 3 code examples: a thorough specification that gives an agent plenty to work with.

With 7,225 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
18/20
Description
15/15
Adoption
16/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 4 days ago, so post-patch-validation is actively maintained.
  • It is released under the CC-BY-SA-4.0 license; check its terms before commercial use.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.

Automated pattern scan on 2026-09-28. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.

post-patch-validation compared with similar skills

All 4 of these similar skills score higher than post-patch-validation; compare them before choosing.

SkillScoreStarsUpdatedFormat
post-patch-validation (this skill)by trailofbits957.2k4d agoSKILL.md
algorithmic-artby anthropics100177.9k5d agoSKILL.md
pptxby anthropics100177.9k5d agoSKILL.md
designby nextlevelbuilder100130.2k6d agoSKILL.md
ui-ux-pro-maxby nextlevelbuilder100130.2k6d agoSKILL.md

Frequently asked questions

How do I install post-patch-validation?
Run npx skills add trailofbits/skills --skill post-patch-validation. The install tabs above show the steps for each supported agent.
Which AI agents does post-patch-validation work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is post-patch-validation safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is CC-BY-SA-4.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is post-patch-validation still maintained?
The repository was last updated 4 days ago, so post-patch-validation is actively maintained.

name: post-patch-validation description: > Validates security patches with reproducible baseline-versus-patched evidence, including original exploits, root-cause variants, behavior preservation, regressions, and newly introduced security failures. Use after a patch exists and before accepting, merging, or reporting it as fixed; also use when an AI-generated patch, remediation commit, pull request, or proposed upstream fix needs adversarial post-patch validation across any language. allowed-tools: Read Write Edit Grep Glob Bash Workflow

Post-Patch Validation

Validate a security patch against the reported bug and the surrounding code it affects. Give the patch author reproducible failures to fix and identify what still needs testing. Apply the same checks to human and agent patches. The diff, author, upstream implementation, and original proof of concept alone cannot establish correctness.

When to Use

  • A security fix, remediation commit, patch file, or pull request already exists.
  • An AI-generated patch needs validation before human review or merge.
  • A fix may cover one exploit path while missing variants of the same root cause.
  • A security fix may alter legitimate behavior or introduce a new vulnerability.
  • A patch author needs concrete failures and coverage gaps before another revision.

When NOT to Use

  • No patch exists yet; use vulnerability discovery or fix implementation first.
  • The task is to review an audit finding against a report without executing patch evidence.
  • The task is only to convert a finding into a permanent project test.
  • The target is remote or production. This skill executes local code and tests only.
  • The user has not authorized execution of the repository's code or test suite.

Quick Start

  1. Pin the vulnerable base and patched input. Prefer immutable commits. For uncommitted work, create a binary patch file first; do not validate in the user's working tree.

  2. Scaffold a pinned plan:

    uv run {baseDir}/scripts/post_patch_validation.py scaffold \
      --repo . \
      --base-ref <vulnerable-ref> \
      --patched-ref <patched-ref> \
      --finding-id <stable-id> \
      --finding-summary "<root cause and impact>" \
      --evidence-level runtime \
      --output post-patch-validation/plan.json
    

    Use --patch-file <path> instead of --patched-ref for a patch artifact. Choose the highest honest evidence level: source for source/patch invariants only, build when target code is compiled or analyzed but the reported behavior is not executed, or runtime when the checks execute the reported behavior and its safety assertions.

  3. Inspect the finding, diff, callers, sibling paths, cleanup/error paths, and existing tests. Populate checks in the generated plan. Run print-schema for the structural schema:

    uv run {baseDir}/scripts/post_patch_validation.py print-schema
    
  4. Run validate-plan for the complete validation, including coverage, command restrictions, and pinned inputs, before executing code:

    uv run {baseDir}/scripts/post_patch_validation.py validate-plan \
      --plan post-patch-validation/plan.json
    
  5. Execute the evidence plan:

    uv run {baseDir}/scripts/post_patch_validation.py run \
      --plan post-patch-validation/plan.json \
      --output post-patch-validation/results
    
  6. Report result.json, report.md, the evidence level, and the complete assessment. Return each finding to the patch author with its check ID, assertion, and saved logs. Identify each validation gap separately, including gaps that coexist with supported findings. After the author revises the patch, pin the new inputs and save a fresh validation run. Preserve the prior evidence. Passing supplied checks still requires human review before acceptance.

Evidence Contract

The runner rejects incomplete plans. Supply at least one check of every kind:

| Kind | Required observation | |---|---| | control | Benign harness succeeds on both base and patch | | exploit | Original safety assertion fails on base and succeeds on patch | | variant | A distinct root-cause variant fails on base and succeeds on patch | | behavior | Unaffected behavior succeeds with byte-identical selected output | | regression | Targeted non-security regression check succeeds on both revisions | | security | Adjacent/new-vulnerability check succeeds on base and patch | | suite | Existing project suite, sanitizer, or deterministic fuzz campaign succeeds on patch |

Commands are argv arrays, never shell strings. Put complex setup in a checked-in or plan artifact script and invoke it with {plan_dir}. The runner fixes locale/timezone/hash-seed inputs, executes checks in lexical ID order, records raw stdout/stderr, and never edits the original worktree. Each check's timeout_seconds defaults to 300 and accepts integers from 1 through 3600. Exceeding the timeout leaves a validation gap. A timeout alone does not establish a regression. Every plan also contains a sorted submodules array ([] when none). Scaffolding infers affected Gitlinks from the changed-file inventory. The runner initializes those pinned commits from the source repository's existing Git module objects, never from .gitmodules network URLs; initialize or fetch them in the source repository before validation.

Exploit and variant checks must prove they ran

A nonzero exit does not mean the vulnerability reproduced. An import error, a failed build, a missing dependency, and a failed safety assertion all exit nonzero and are indistinguishable to the runner. Every exploit and variant check must print and flush PPV_REACHED immediately before it evaluates its assertion, on both revisions:

"argv": ["python3", "-c", "import app; value = app.render('<'); print('PPV_REACHED', flush=True); assert value == '&lt;'"]

The token is also in the environment as PPV_REACHED_MARKER. It must land on stdout, as a line of its own. Stderr is not scanned, because a Python SyntaxError traceback echoes the offending source and would otherwise satisfy the check for a harness that executed nothing. A run without it leaves a marker_missing gap. Independent findings from other checks remain in the result. Flush explicitly: a harness whose payload segfaults or calls _exit loses buffered output and forfeits its own evidence.

These checks also run side-blind. {side} is not expanded for them, PPV_SIDE is absent from their environment, the checkout directory is randomly named, and the plan validator rejects any exploit or variant check whose argv or env mentions either. An assertion that can see which revision it is on can assert on that instead of on the code, which is the cheapest possible way to fake a reproduction followed by a fix.

Environment

Checks run under a fixed minimal environment: PATH, HOME, and a handful of temp/user keys, plus LANG/LC_ALL=C, TZ=UTC, PYTHONHASHSEED=0, NO_COLOR, TERM=dumb. Everything else in the caller's environment is dropped. Toolchains that need more get it explicitly:

uv run {baseDir}/scripts/post_patch_validation.py run \
  --plan post-patch-validation/plan.json \
  --output post-patch-validation/results \
  --allow-env JAVA_HOME --allow-env CARGO_HOME

Forwarded names and values are recorded in result.json. A requested variable that is unset is an error, not an empty string. Two classes are refused outright: names that read as credentials (*SECRET*, *TOKEN*, *API_KEY*, …), because the value would be written into the result; and names that change what executes (LD_PRELOAD, BASH_ENV, NODE_OPTIONS, GIT_SSH_COMMAND, …), because those variables can change which code executes. The runner's fixed variables and every PPV_* name are also reserved and cannot be forwarded.

Placeholders expanded in argv and per-check env values: {checkout} (the revision under test), {plan_dir} (an isolated copy of the plan artifacts for that one invocation), {scratch} (a fresh opaque directory for that one check invocation), and {side} (base or patched, and not available to exploit/variant checks). The same values arrive as PPV_CHECKOUT, PPV_PLAN_DIR, PPV_SCRATCH, PPV_SIDE, and PPV_CASE_ID. Write only under {scratch}; the evidence directory path is not passed to checks. Base and patched invocations do not share runner-managed scratch, plan, or worktree roots. After each invocation exits, its scratch tree is archived under the deterministic results/scratch/<check-id-and-side> path, its private plan copy is discarded, and every readable argv element that resolves to a file is hashed in argv_files. Files inside the isolated plan or checkout roots are additionally retained under results/helpers/<sha256> up to 16 MiB; the record explains why any other file was not archived. Use a dedicated directory for plan.json: its sibling files and directories are copied into each invocation's {plan_dir}. Keep helper code under that directory's checks/ directory or checked into the target repository so its bytes are reviewable. The machine plan containing commit pins, the current output directory, and detected prior result trees are excluded; symlinks are rejected. The clean snapshot remains only in runner memory, and exploit/variant sides execute in random order while evidence filenames remain deterministic. Stdout/stderr use anonymous or randomly named capture descriptors and are copied to the named evidence files only after the child exits, so fd inspection cannot disclose the side label.

This isolation is not a host sandbox: checks run with the caller's privileges and a malicious helper could use arbitrary external state or deliberately infer the revision from source or Git metadata. Inspect the content-addressed helper artifacts, and use an OS/container sandbox when the check code itself is untrusted.

Active validation worktrees are Git-locked with random owner tokens backed by kernel file locks, so another concurrent validator cannot prune them and PID reuse cannot impersonate an owner. If the runner is forcibly killed, the next run unlocks stale validator-owned registrations. For manual recovery, inspect git worktree list, then use git worktree unlock <path> and git worktree remove --force <path> (or git worktree prune after the path is gone).

Read evidence-model.md when designing coverage, selecting variants, or interpreting findings and validation gaps. Do not read it for routine CLI execution.

Coverage Rules

  • Derive variants from the root cause, not cosmetic mutations of the original payload.
  • Enumerate sibling call sites, alternate callbacks/outputs, error paths, teardown, ownership, serialization, and boundary values touched by the fix.
  • Make each exploit or variant test assert the safe behavior. It must fail on the vulnerable base; a test that passes on both revisions proves nothing about remediation. Read the base-side stderr and confirm the failure is the assertion you wrote, not a harness that never got there.
  • Keep exploit and variant assertions limited to the security invariant. Test liveness, exact error types/messages, timing, and compatibility separately as behavior or regression checks; otherwise an unrelated contract change can masquerade as proof that the vulnerability remains.
  • Keep the control harness benign and make it exercise the changed component. It establishes that the harness works on both revisions. Failed controls leave gaps and prevent attributing other failures to the patch. The raw observations remain available for review.
  • Use behavior only for behavior that should remain unchanged. Exact output comparison is deliberate; move unstable values behind a deterministic test harness instead of normalizing them away in prose.
  • Make security chec

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars7.2k
CategorySecurity
Updated4d ago
Forks615

Languages

Python

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions