SkillAgentSearch skills...

skill-auditor

A comprehensive auditor for any agent skill — including Manus, OpenClaw/ClawHub, Claude, LobeHub, or custom SKILL.md-based skills. Use this skill whenever a user wants to evaluate, audit, review, score, or quality-check an agent skill before publishing, updating, or deploying.

Install / Use

npx skills add aipoch/medical-research-skills --skill skill-auditor

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

94/100

Category

Operations

Supported Platforms

Claude Code
Zed

Our assessment of skill-auditor

skill-auditor scores 94/100 on our quality scale, 132nd of 748 Operations skills we index (top 18%).

Its SKILL.md is 27 KB long, well organised into 38 sections with 13 code examples: a thorough specification that gives an agent plenty to work with.

With 1,916 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
14/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 15 days ago, so skill-auditor is actively maintained.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

skill-auditor compared with similar skills

All 4 of these similar skills score higher than skill-auditor; compare them before choosing.

SkillScoreStarsUpdatedFormat
skill-auditor (this skill)by aipoch941.9k15d agoSKILL.md
algorithmic-artby anthropics100177.9k10d agoSKILL.md
pptxby anthropics100177.9k10d agoSKILL.md
designby nextlevelbuilder100130.2k11d agoSKILL.md
ui-ux-pro-maxby nextlevelbuilder100130.2k11d agoSKILL.md

Frequently asked questions

How do I install skill-auditor?
Run npx skills add aipoch/medical-research-skills --skill skill-auditor. The install tabs above show the steps for each supported agent.
Which AI agents does skill-auditor work with?
It is written for Claude Code and Zed, as a SKILL.md file. Other agents that read the same format can often use it too.
Is skill-auditor safe to use?
It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is skill-auditor still maintained?
The repository was last updated 15 days ago, so skill-auditor is actively maintained.

name: skill-auditor description: A comprehensive auditor for any agent skill — including Manus, OpenClaw/ClawHub, Claude, LobeHub, or custom SKILL.md-based skills. Use this skill whenever a user wants to evaluate, audit, review, score, or quality-check an agent skill before publishing, updating, or deploying. Covers two hard veto gates (structural redlines + research integrity redlines), static quality scoring across 25 criteria (ISO 25010 + OpenSSF + Agent), dynamic test input generation, multi-mode execution testing, multi-layer output evaluation with five specialized category rubrics (Evidence Insight / Protocol Design / Data Analysis / Academic Writing / Other), a Research Veto that applies to all four research categories, human eval viewer generation, actionable P0/P1/P2 optimization recommendations, and automatic skill improvement that outputs a polished, production-ready SKILL.md. Also use whenever a user says "audit my skill", "evaluate my skill", "improve my skill", or wants a corrected version after evaluation. license: MIT skill-author: AIPOCH

Skill Auditor

This skill provides a standardized, end-to-end process for auditing any agent skill — from structural integrity to live functional performance. It combines two independent veto gates, static analysis across 25 criteria, dynamic execution, five-category specialized scoring, and human review into a single coherent workflow.

Audit Pipeline

Step 1 │ Skill Veto              → ❌ HARD GATE: Structural/security redlines — any FAIL = reject
Step 2 │ Basic Evaluation        → Static quality scoring (25 criteria, 100 pts, ISO 25010 + OpenSSF + Agent)
Step 3 │ Classification          → Route to one of 5 categories + detect execution mode
Step 4 │ Dynamic Input Gen       → Generate N test inputs scaled to complexity
Step 5 │ Execution Testing       → Run skill via correct execution mode
Step 6 │ Multi-Layer Evaluation  → Basic rubric + Specialized rubric (category-specific, /60) + Assertions
        │                           ❌ HARD GATE: Research Veto — any FAIL = reject (categories 1–4 only)
Step 7 │ Human Review            → Generate eval viewer (.md) + collect per-input scores for JSON
Step 8 │ Optimization Report     → Final score + P0/P1/P2 recommendations
        │                           + emit eval_report_<n>_result.json for frontend visualization

Two hard rejection gates run at Steps 1 and 6. Both are mandatory and cannot be skipped. All other steps run sequentially. Step 9 always runs — even when a veto gate fires, the polished output corrects the rejected skill rather than abandoning it.


Language Policy

All audit output must be written in English, regardless of the language used in the user's request or in the submitted skill.

This applies to every artifact produced by this skill:

  • Veto reports (Step 1 and Step 6)
  • Static evaluation scores and notes (Step 2)
  • Generated test inputs (Step 4)
  • Execution summaries and per-output evaluations (Steps 5–6)
  • The eval viewer .md file (Step 7)
  • The final optimization report and JSON (Step 8)

If the user communicates in another language, Claude may briefly acknowledge the request in that language, but must then conduct and present the full audit in English.


Step 1: Skill Veto — Structural Redlines ❌

HARD GATE. Any FAIL = immediate rejection. Do not proceed to Step 2.

Read the target skill's SKILL.md and any bundled scripts. Check all four dimensions:

→ Full criteria: references/basic_veto.md

| Dimension | Immediate Rejection Triggers | |---|---| | T1. Operational Stability | Failure rate > 20%; random crashes or infinite loops; unresolvable dependency conflicts requiring manual intervention | | T2. Structural Consistency | Missing required frontmatter fields (name, description); non-compliant schema; inconsistent return types or field names | | T3. Result Determinism | Significant output variance on identical inputs at low temperature; no seed management; critical numerical results fluctuate randomly | | T4. System Security | Direct execution of raw user-provided strings (eval/exec); no input filtering; prompt injection vectors present in scripts or instructions |

► If any T1–T4 dimension is FAIL: stop immediately. Output the rejection report below and do not continue.

SKILL VETO — REJECTED
══════════════════════════════════
Skill: <n>
Reason: Failed structural redline check
T1. Stability    : PASS / FAIL — <reason>
T2. Contract     : PASS / FAIL — <reason>
T3. Determinism  : PASS / FAIL — <reason>
T4. Security     : PASS / FAIL — <reason>

This skill must not be deployed. Fix all FAIL dimensions before resubmitting.
══════════════════════════════════

Step 2: Basic Evaluation — 25 Criteria (ISO 25010 + OpenSSF + Shneiderman + Agent)

Read the skill's full SKILL.md and all bundled files. Score each of the 25 criteria from 0–4.

→ Full rubric with per-level descriptions: references/basic_evaluation.md

| # | Category (Framework) | Criteria | Max | |---|---|---|---| | 1 | Functional Suitability (ISO 25010) | Completeness, Correctness, Appropriateness | 12 | | 2 | Reliability (ISO 25010) | Fault Tolerance, Error Reporting, Recoverability | 12 | | 3 | Performance & Context (ISO 25010 + Agent) | Token Cost, Execution Efficiency | 8 | | 4 | Agent Usability (Shneiderman · Gerhardt-Powals) | Learnability, Consistency, Feedback Design, Error Prevention | 16 | | 5 | Human Usability (Tognazzini · Norman) | Discoverability, Forgiveness | 8 | | 6 | Security (ISO 25010 + OpenSSF) | Credential Safety, Input Validation, Data Safety | 12 | | 7 | Maintainability (ISO 25010) | Modularity, Modifiability, Testability | 12 | | 8 | Agent-Specific (Novel) | Trigger Precision, Progressive Disclosure, Composability, Idempotency, Escape Hatches | 20 |

Basic subtotal: __ / 100


Step 3: Classification + Execution Mode Detection

3.1 Classify the Skill

Read the skill's description frontmatter and ## When to Use section.

→ Full category definitions: references/classification.md

| # | Category | Typical Skills | |---|---|---| | 1 | Evidence Insight | Search strategy builders, database scouts, critical appraisal tools, evidence synthesizers | | 2 | Protocol Design | Experimental design generators, study-type advisors, statistical power planners, validation strategists | | 3 | Data Analysis | R/Python code generators, bioinformatics pipelines, statistical modeling tools, ML workflows | | 4 | Academic Writing | SCI manuscript writers, abstract generators, methods/discussion drafters, cover-letter tools | | 5 | Other (General / Non-Research) | All skills that do not fall into categories 1–4 |

Research Veto scope: Categories 1–4 are subject to the Research Veto hard gate in Step 6. Category 5 is exempt.

3.2 Detect Execution Mode

Inspect the skill to determine how it is meant to be invoked.

| Mode | Indicators | How to Run in Step 5 | |---|---|---| | A: Direct | Only SKILL.md instructions, no scripts | Follow SKILL.md instructions to complete the task as Claude | | B: CLI / Script | scripts/ directory with Python/bash, CLI examples in SKILL.md | Execute via bash: python scripts/xxx.py <args> | | C: API | API endpoint patterns, fetch/curl usage in SKILL.md | Simulate or call the API as documented | | D: Hybrid | Both instructions and scripts/API | Run script for deterministic parts; Claude for reasoning/generation parts |

Record the detected mode. It will be used in Step 5.


Step 4: Dynamic Input Generation

Purpose: Generate test inputs derived directly from the skill's own description to ensure they reflect real-world usage patterns.

4.1 Assess Skill Complexity

| Complexity Level | Criteria | Test Input Count (N) | |---|---|---| | Simple | Single task type, narrow scope, < 3 reference files, no branching workflow | 3 inputs | | Moderate | 2–3 task types, some branching, 3–5 reference files, moderate scope | 5 inputs | | Complex | Multiple task types, branching logic, 5+ reference files, broad or specialized scope | 7 inputs |

Declare: Complexity: [Simple / Moderate / Complex] → Generating N inputs

4.2 Generate N Test Inputs

Use this distribution based on N:

| Slot | Type | Always Include? | |---|---|---| | Input 1 | Canonical / happy path | ✅ Always | | Input 2 | Variant A (different valid use case) | ✅ Always | | Input 3 | Edge / boundary | ✅ Always | | Input 4 | Variant B (third central use case) | If N ≥ 5 | | Input 5 | Stress / complex / multi-part | If N ≥ 5 | | Input 6 | Scope boundary (slightly outside) | If N = 7 | | Input 7 | Adversarial / ambiguous | If N = 7 |

Format each as a realistic user message. Do not include expected answers.

Output:

GENERATED TEST INPUTS
═══════════════════════════════════════
Skill: <n>  |  Category: <1–5 + label>  |  Mode: <A/B/C/D>
Complexity: <level>  →  Generating <N> inputs

Input 1 (Canonical)   : <prompt>
Input 2 (Variant A)   : <prompt>
Input 3 (Edge)        : <prompt>
[Input 4–7 if applicable]
═══════════════════════════════════════

Step 5: Execution Testing

Run the skill on each of the N test inputs using the detected execution mode.

Mode A: Direct Execution

Load the skill's SKILL.md. Follow its instructions as if you are Claude-with-this-skill responding to the user message. Complete the task in full.

Mode B: CLI / Script Execution

python scripts/<script_name>.py "<input_text_or_path>"

Capture stdout/stderr. If execution fails, record the error and continue.

Mode C: API Execution

Follow the API usage pattern documented in the skill. Construct the request, execute, capture response. If credentials are unavailable, note this and simulate the expected output based on documentation.

Mode D: Hybrid Execution

Run the script/API component first. Pass its output to Claude for the reasoning/generation component.

Execution Log

For each input:

─── Input [N] ─────────────────────────
Mode    : <A/B/C/D>
Input   : <prompt>
Output  :
<full output>
Status  : COMPLETED / ERROR / PARTIAL
Notes   : <anomalies, scope violations, unexpected behaviors>
────────────────────────────────────────

Step 6: Multi-Layer Output Evaluation

Evaluate all N outputs across three parallel layers.

Layer 1: Basic Rubric Scoring

→ Full criteria: references/basic_evaluation.md

For each output, score four aggregate dimensions (0–10 each):

  • Functional Correctness (did output complete the task correctly and fully?): /10
  • Reliability & Clarity (well-structured, consistent, clear feedback?): /10
  • Efficiency (concise, no padding, no unnecessary context bloat?): /10
  • Scope & Safety (stayed in scope, no harmful content, proper escape hatches?): /10

Per-output basic score: /40

Layer 2: Specialized Rubric Scoring

Apply the rubric corresponding to the category from Step 3:

| Category | Reference File | Max | |---|---|---| | 1 — Evidence Insight | references/specialized_evaluation_literature.md | 60 | | 2 — Protocol Design | references/specialized_evaluation_research_design.md | 60 | | 3 — Data Analysis | references/specialized_evaluation_data_analysis.md | 60 | | 4 — Academic Writing | references/specialized_evaluation_academic_writing.md | 60 | | 5 — Other | references/specialized_evaluation_other.md | 60 |

Per-output specialized score: /60

Layer 3: Assertion Checks

For each output, write and eva

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars1.9k
CategoryOperations
Updated15d ago
Forks175

Languages

Python

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions