6 skills found
Evol-ai / SkillCompassEvaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.
EternalWavee / benchmark-research-skillClaude Code skill for benchmark research. Survey papers to find datasets, metrics, and evaluation protocols used in a research direction.
opendatalab / omnidocbench-eval-helperHelp users deploy, validate, run, and parse OmniDocBench evaluations. Use this skill whenever the user mentions OmniDocBench, document parsing/OCR benchmark scoring, MinerU or other model evaluation on OmniDocBench, CDM formula metrics, end2end/md2md configs, Docker/conda deployment, remote SSH/H-cl…
actboy168 / luamakeLuamake 构建系统指南——用于当前项目的 `luamake` / `make.lua` / Ninja 生成流程。当用户需要编写、修改、排查或理解 `make.lua`、目标定义、`lm:conf`、`deps` / `objdeps`、代码生成、Lua C 模块、Bee 运行时集成,或需要解决当前项目中由 `luamake` 驱动的构建问题时,使用此 skill。即使用户没有明确提到 `luamake`,但上下文明显是在处理本项目的构建脚本、构建目录、目标依赖或 `luamake` 生成的 Ninja 流程,也应使用此 skill。不要把它用于纯通用的 C/C++ 编译知识、与本项目无…
microsoft-foundry / azure-ai-fine-tuningUse when the user wants to fine-tune a model on Azure AI Foundry — including dataset preparation, training, evaluation, and deployment.
rwi001 / llm-evaluationImplement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking