8 skills found
google-gemini / behavioral-evalsGuidance for creating, running, fixing, and promoting behavioral evaluations
refreshdotdev / web-eval-agentAn MCP server that autonomously evaluates web applications.
opendatalab / omnidocbench-eval-helperHelp users deploy, validate, run, and parse OmniDocBench evaluations. Use this skill whenever the user mentions OmniDocBench, document parsing/OCR benchmark scoring, MinerU or other model evaluation on OmniDocBench, CDM formula metrics, end2end/md2md configs, Docker/conda deployment, remote SSH/H-cl…
PAIR-code / eval-prSet up a git worktree to run and evaluate an upstream Pull Request locally in a flat worktree bare clone repository layout.
evalstate / fast-agentCode, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support
majiayu000 / issue-size-estimationIssue粒度判定とコード行数見積もりの基準、サイズラベル、分割判定ロジックを定義
greynewell / mcpbrBenchmark your MCP server.
Evol-ai / SkillCompassEvaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.