9 skills found
openclaw / canvasPresent HTML on connected OpenClaw node canvases, navigate/eval/snapshot, and debug canvas host URLs.
Evol-ai / SkillCompassEvaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.
devallibus / shiplogSUPERCHARGE AI-assisted development by using Git. Cross-model review gates, evidence-linked closure, verification profiles, model-tier routing, artifact envelopes, and provenance signing — all from a single skill for Claude Code, Codex, and Cursor.
EternalWavee / benchmark-research-skillClaude Code skill for benchmark research. Survey papers to find datasets, metrics, and evaluation protocols used in a research direction.
lee-fuhr / claude-session-indexIndex, search, and analyze your Claude Code sessions. Full-text search, conversation retrieval, analytics, and cross-session synthesis.
Agents365-ai / scholar-deep-research8-phase literature-review pipeline. 7 federated sources, dedup, ranked retrieval, citation chasing, self-critique, 5 report archetypes.
opendatalab / omnidocbench-eval-helperHelp users deploy, validate, run, and parse OmniDocBench evaluations. Use this skill whenever the user mentions OmniDocBench, document parsing/OCR benchmark scoring, MinerU or other model evaluation on OmniDocBench, CDM formula metrics, end2end/md2md configs, Docker/conda deployment, remote SSH/H-cl…
microsoft-foundry / azure-ai-fine-tuningUse when the user wants to fine-tune a model on Azure AI Foundry — including dataset preparation, training, evaluation, and deployment.
rwi001 / llm-evaluationImplement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking