8 skills found
google-gemini / behavioral-evalsGuidance for creating, running, fixing, and promoting behavioral evaluations
PatrickJS / google-adkGoogle Agent Development Kit rules for agents, tools, sessions, memory, artifacts, evaluation, and deployment
MCPJam / inspectorTesting and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
trpc-group / trpc-agent-goA Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
EternalWavee / benchmark-research-skillClaude Code skill for benchmark research. Survey papers to find datasets, metrics, and evaluation protocols used in a research direction.
opendatalab / omnidocbench-eval-helperHelp users deploy, validate, run, and parse OmniDocBench evaluations. Use this skill whenever the user mentions OmniDocBench, document parsing/OCR benchmark scoring, MinerU or other model evaluation on OmniDocBench, CDM formula metrics, end2end/md2md configs, Docker/conda deployment, remote SSH/H-cl…
microsoft-foundry / azure-ai-fine-tuningUse when the user wants to fine-tune a model on Azure AI Foundry — including dataset preparation, training, evaluation, and deployment.
rwi001 / llm-evaluationImplement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking