18 skills found
refreshdotdev / web-eval-agentAn MCP server that autonomously evaluates web applications.
PAIR-code / eval-prSet up a git worktree to run and evaluate an upstream Pull Request locally in a flat worktree bare clone repository layout.
opendatalab / omnidocbench-eval-helperHelp users deploy, validate, run, and parse OmniDocBench evaluations. Use this skill whenever the user mentions OmniDocBench, document parsing/OCR benchmark scoring, MinerU or other model evaluation on OmniDocBench, CDM formula metrics, end2end/md2md configs, Docker/conda deployment, remote SSH/H-cl…
google-gemini / behavioral-evalsGuidance for creating, running, fixing, and promoting behavioral evaluations
yllibed / replOne .NET command graph. CLI, REPL, remote sessions, MCP.
rwi001 / llm-evaluationImplement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking
evalstate / fast-agentCode, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support
char5742 / 00_basicユーザーはRooよりプログラミングが得意ですが、時短のためにRooにコーディングを依頼しています。 2回以上連続でテストを失敗した時は、現在の状況を整理して、一緒に解決方法を考えます。 私は GitHub
StephanSchmidt / loupeOpen-source CLI that measures AI coding-assistant impact (Claude Code, Copilot, Cursor, Aider) across Bitbucket+Jira, GitHub, GitLab, Azure DevOps, or Linear — and renders a reveal.js exec deck. Runs locally, no SaaS, no data leaves your environment.
depwire / depwireThe missing context layer for AI-assisted refactoring
MCPJam / inspectorTesting and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
gsd-build / get-shit-doneA light-weight and powerful meta-prompting, context engineering and spec-driven development system for Claude Code by TÂCHES.
kyouyap / 00_basicユーザーはRooよりプログラミングが得意ですが、時短のためにRooにコーディングを依頼しています。 2回以上連続でテストを失敗した時は、現在の状況を整理して、一緒に解決方法を考えます。 私は GitHub
sbhooley / ainativelangAINL helps turn AI from "a smart conversation" into "a structured worker." It is designed for teams building AI workflows that need multiple steps, state and memory, tool use, repeatable execution, validation and control, and lower dependence on long prompt loops.
Evol-ai / SkillCompassEvaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.
oboard / claude-code-revRunnable ClaudeCode source code
Ecolash / MCPlex-AI-v1.0𝗠𝗼𝗱𝗲𝗹 𝗖𝗼𝗻𝘁𝗲𝘅𝘁 𝗣𝗿𝗼𝘁𝗼𝗰𝗼𝗹 (𝗠𝗖𝗣) 𝗕𝗮𝘀𝗲𝗱 𝗖𝗟𝗜 𝗔𝗜 | 𝗧𝗼𝗼𝗹 𝗜𝗻𝘁𝗲𝗴𝗿𝗮𝘁𝗶𝗼𝗻 | 𝗚𝗲𝗺𝗶𝗻𝗶 𝟮.𝟬
adityaarsharma / pickleFree, local, open-source MCP server that audits your ClickUp, Slack & Teams for what fell through the cracks — stale tasks, dropped promises, decisions lost in DMs. No account, no telemetry. Works in Claude, Cursor, Codex, Cline, Zed.