SkillAgentSearch skills...

write-code-eval

Write code evaluators for known failure modes with objective rules

Install / Use

npx skills add ai-evals-course/evals-skills --skill write-code-eval

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

64/100

Supported Platforms

Universal

Tags

Our assessment of write-code-eval

write-code-eval scores 64/100 on our quality scale, 3532nd of 4,259 Development & Engineering skills we index.

Its SKILL.md is 1.5 KB long, lightly structured (2 headings) and no code examples: moderately detailed.

With 1,267 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
20/30
Structure
5/20
Description
12/15
Adoption
13/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 15 days ago, so write-code-eval is actively maintained.
  • No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
  • Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

write-code-eval compared with similar skills

All 4 of these similar skills score higher than write-code-eval; compare them before choosing.

SkillScoreStarsUpdatedFormat
write-code-eval (this skill)by ai-evals-course641.3k15d agoSKILL.md
ai-job-searchby MadsLorentzen10044.6k1d agoCLAUDE.md
claude-howtoby luongnv8910041.7ktodayCLAUDE.md
algorithmic-artby anthropics100177.9k8d agoSKILL.md
pptxby anthropics100177.9k8d agoSKILL.md

Frequently asked questions

How do I install write-code-eval?
Run npx skills add ai-evals-course/evals-skills --skill write-code-eval. The install tabs above show the steps for each supported agent.
Which AI agents does write-code-eval work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is write-code-eval safe to use?
It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is write-code-eval still maintained?
The repository was last updated 15 days ago, so write-code-eval is actively maintained.

name: write-code-eval description: > Write code evaluators for known failure modes with objective rules. Use when code can check the rule from a trace, with or without a reference answer. Use write-judge-prompt when the rule requires interpretation.

Write a code evaluator

Start with a failure mode found through error analysis. Write one check for that failure mode, much like a unit test that asserts what should hold for each trace.

  1. State the rule and identify the trace fields or reference data the check needs. If the rule requires interpretation, use write-judge-prompt.
  2. Implement the check in the project's language and eval framework. Return a result and a reason in the format that framework expects.
  3. Test known passes and failures, including borderline cases. Run the check on available traces and inspect mistakes. If the rule uses a proxy for human judgment, compare its results with human labels.

Examples

| Failure mode | Possible check | |---|---| | Invalid output structure | Parse the output and check required fields | | Missing or forbidden text | Match a string or pattern | | Citation not in retrieved documents | Compare cited IDs with retrieved IDs | | Bad tool call | Check arguments against the tool schema or run the call in a safe test environment | | Wrong value | Compare the output with a reference value |

Choose the check from the failure rule. For a failure with both objective and interpretive parts, check the objective part with code and use a judge for the rest.

Related Skills

View on GitHub
GitHub Stars1.3k
CategoryDevelopment
Updated15d ago
Forks88

Trust signals

88/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 medium