21 skills found
google-gemini / behavioral-evalsGuidance for creating, running, fixing, and promoting behavioral evaluations
santifer / career-opsOpen-source AI job search: scan job portals, evaluate listings with a structured A-F rubric into a 1.0-5.0 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Code, Codex, OpenCode, Antigravity…)
PatrickJS / google-adkGoogle Agent Development Kit rules for agents, tools, sessions, memory, artifacts, evaluation, and deployment
MadsLorentzen / ai-job-searchThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
decolua / 9routerUnlimited FREE AI coding. Connect Claude Code, Codex, Cursor, Cline, Copilot, Antigravity to FREE Claude/GPT/Gemini via 40+ providers. Auto-fallback, RTK -40% tokens, never hit limits.
google / adk-goAn open-source, code-first Go toolkit for building, evaluating, and deploying sophisticated AI agents with flexibility and control.
cporter202 / API-mega-listThis GitHub repo is a powerhouse collection of APIs you can start using immediately to build everything from simple automations to full-scale applications. One of the most valuable API lists on GitHub—period. 💪
Kiln-AI / KilnBuild, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
evalstate / fast-agentCode, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support
MCPJam / inspectorTesting and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
IBM / AssetOpsBenchAssetOpsBench - Industry 4.0: A unified benchmark and framework for building, orchestrating, and evaluating domain-specific AI agents for Industry 4.0 asset operations and maintenance, with 460+ scenarios, 5 specialist agents (IoT, FMSR, TSFM, Work Order,...), and multi-agent orchestration blueprint…
trpc-group / trpc-agent-goA Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
refreshdotdev / web-eval-agentAn MCP server that autonomously evaluates web applications.
Evol-ai / SkillCompassEvaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.
EternalWavee / benchmark-research-skillClaude Code skill for benchmark research. Survey papers to find datasets, metrics, and evaluation protocols used in a research direction.
opendatalab / omnidocbench-eval-helperHelp users deploy, validate, run, and parse OmniDocBench evaluations. Use this skill whenever the user mentions OmniDocBench, document parsing/OCR benchmark scoring, MinerU or other model evaluation on OmniDocBench, CDM formula metrics, end2end/md2md configs, Docker/conda deployment, remote SSH/H-cl…
actboy168 / luamakeLuamake 构建系统指南——用于当前项目的 `luamake` / `make.lua` / Ninja 生成流程。当用户需要编写、修改、排查或理解 `make.lua`、目标定义、`lm:conf`、`deps` / `objdeps`、代码生成、Lua C 模块、Bee 运行时集成,或需要解决当前项目中由 `luamake` 驱动的构建问题时,使用此 skill。即使用户没有明确提到 `luamake`,但上下文明显是在处理本项目的构建脚本、构建目录、目标依赖或 `luamake` 生成的 Ninja 流程,也应使用此 skill。不要把它用于纯通用的 C/C++ 编译知识、与本项目无…
microsoft-foundry / azure-ai-fine-tuningUse when the user wants to fine-tune a model on Azure AI Foundry — including dataset preparation, training, evaluation, and deployment.
rwi001 / llm-evaluationImplement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking
Duds / accessibilityAn MCP for testing, evaluating and auditing website accessibility.
luanmorenommaciel / project-recapGenerate a visual HTML project recap — rebuild mental model of a project's current state, recent decisions, and cognitive debt hotspots