gpu-perf-tune
Agent Skills and an MCP server for GPU performance profiling, benchmarking, optimization, and reporting, with an inference focus.
Install / Use
claude mcp add cfregly -- npx -y github:cfregly/gpu-perf-tuneIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AI & Machine LearningSupported Platforms
Tags
Our assessment of gpu-perf-tune
gpu-perf-tune scores 81/100 on our quality scale, 734th of 969 AI & Machine Learning skills we index.
Its MCP Server is 9.4 KB long, well organised into 20 sections with 4 code examples: a thorough specification that gives an agent plenty to work with.
It has 3 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated yesterday, so gpu-perf-tune is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 92/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
gpu-perf-tune compared with similar skills
All 4 of these similar skills score higher than gpu-perf-tune; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| gpu-perf-tune (this skill)by cfregly | 81 | 3 | 1d ago | MCP Server |
| claude-memby thedotmack | 100 | 97.5k | today | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 93.0k | 21d ago | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.5k | 1d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.6k | today | CLAUDE.md |
Frequently asked questions
- How do I install gpu-perf-tune?
- Run
claude mcp add cfregly -- npx -y github:cfregly/gpu-perf-tune. The install tabs above show the steps for each supported agent. - Which AI agents does gpu-perf-tune work with?
- It is written for Claude Code, Claude Desktop and OpenAI Codex, as a MCP Server file. Other agents that read the same format can often use it too.
- Is gpu-perf-tune safe to use?
- It is MIT-licensed and scores 92/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is gpu-perf-tune still maintained?
- The repository was last updated yesterday, so gpu-perf-tune is actively maintained.
Skill content
View source on GitHubgpu-perf-tune
GPU performance engineering for agents and MCP clients, with an inference-focused skills package and reusable cluster utilities. The project ships 32 task-oriented Agent Skills, a bundled MCP server, safety guards, workload proof schemas, and local validation tools. Together they cover first estimates, benchmark sweeps, kernel profiling, speed-of-light analysis, optimization, and evidence-backed reports.
The work comes from GPU fleet performance practice. Every result stays a candidate until the workload, baseline, hardware, precision, parallelism, and engine version are recorded and the result survives a skeptical check.
Names and boundaries
| Name | Meaning |
| --- | --- |
| gpu-perf-tune | The project and repository. It owns the skills, MCP server, guards, schemas, examples, and validation. |
| profile-and-optimize | The skills package under plugins/profile-and-optimize/. Claude Code consumes it as a plugin. Codex installs the same skills from a clone. |
| profile_and_optimize | The configured MCP server key. It is usable without the Claude Code plugin. |
| profile_and_optimize_mcp | The Python module that serves the MCP tools and resources. |
Claude Code is one supported client, not the project boundary. Client support depends on the surface being installed.
Client support
| Client | Agent Skills | MCP server | Safety guards |
| --- | --- | --- | --- |
| Claude Code | First-class marketplace plugin | Plugin or repo installer | Provenance hook is installed. Enforcement is opt-in |
| Codex | First-class repo installer into ~/.agents/skills | Repo installer | No client hook adapter. MCP acknowledgement gates still apply |
| Cursor, Gemini CLI, and Antigravity | Best-effort helpers | Best-effort repo installer | No release-tested adapter |
| Other stdio MCP clients | Client-owned discovery | Documented stdio command | No packaged adapter |
The SKILL.md sources follow the open Agent Skills standard. This repository
maintains and tests the Claude Code and Codex paths. The MCP protocol remains
client-neutral. Configuration helpers for other clients are available, but
they are not part of the release bar.
What it covers
- Estimate, benchmark, and sweep:
inference-performance-hintsfor rough performance bounds,inference-perf-benchfor load sweeps,inference-tune-sweepfor engine knobs,inference-model-evalfor quality gates, andperf-baseline-recordwithperf-baseline-difffor regressions. - Profile:
inference-workload-profile,inference-kernel-profilefor nsys,inference-kernel-ncu-profilefor per-kernel roofline work,inference-dcgm-correlate,analyze-zymtrace-workload,inference-graph-diff, andmirage-graph-coverage. - Optimize:
inference-model-optimize,inference-quantize-calibrate, the speculative decode train, tune, and service skills,inference-decode-step-budget,inference-capacity-sizing, andinference-known-good-config. - Report and track:
inference-perf-tune-report,inference-perf-synthesize,inference-fleet-leaderboard,inference-value-ledger,evidence-bundle-init, and the anchored Prometheus and zymtrace query skills.
The documented command-line path remains available when an external observability MCP server is absent.
Quickstart
Inspect the project without a GPU
git clone https://github.com/cfregly/gpu-perf-tune.git
cd gpu-perf-tune
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements/check.txt
make demo
make check
make workload-proof-check
make demo prints the skill and MCP tool surface. A real performance run needs
the bundled server, the target workload, and suitable GPU hardware.
Codex: skills and MCP
From a clone of this repository:
make install-skills CLIENT=codex
make install-mcp CLIENT=codex
codex mcp get profile_and_optimize
The registration check should report enabled: true. Run
make smoke-mcp-runtime for a live server handshake, then restart Codex after
installation. Codex discovers the linked skills under
~/.agents/skills. Use /skills, invoke a skill with $skill-name, or
describe the task and let Codex select one. The MCP configuration is shared by
Codex CLI, the IDE extension, and the desktop app. See the official Codex
Skills and
MCP references for the client-side
contracts.
Claude Code: skills, MCP, and optional provenance enforcement
claude plugin marketplace add cfregly/gpu-perf-tune
claude plugin install --scope user \
profile-and-optimize@profile-and-optimize-plugins
# Install the bundled MCP server inside the current plugin cache entry.
# Add --full when you need the PDF report dependencies.
bash "$(ls -dt ~/.claude/plugins/cache/profile-and-optimize-plugins/profile-and-optimize/*/server/install.sh | head -1)"
claude mcp get plugin:profile-and-optimize:profile_and_optimize
The health check should report Status: ✔ Connected. Restart Claude Code.
Invoke a skill such as /inference-perf-bench, or describe the task and let
Claude Code select a matching skill. The plugin installs its provenance hook,
but the hook remains inactive until
PROVENANCE_COMMIT_GATE=ask or PROVENANCE_COMMIT_GATE=deny is present in the
Claude hook environment. Install jq on PATH before enabling either mode.
Other clients: best-effort helpers
The repository keeps configuration helpers for Cursor, Gemini CLI, and Google Antigravity. They are useful starting points, but changes to those clients do not block a release. The full command list and generic stdio form live in the MCP installation reference.
make install-skills CLIENT=cursor
make install-mcp CLIENT=cursor
Upgrading
Read docs/UPGRADING.md before changing release lines. It
lists the client refresh commands and required installer, AI tuning, MLPerf
rules, and MCP acknowledgement changes.
Value bar
Every benchmark result, optimization claim, and generated report starts as a candidate. It must be adversarially-confirmed to add value before it ships. The workload is named, the baseline is fair, a skeptic has tried to break the finding, and the receipt maps to lower cost, faster runtime, higher throughput, better reliability, or a clearer operator action.
Workload proof contract
docs/workload-proof-packet.md defines the
GPU inference packet shape for neocloud buyers and workflow handoffs.
make workload-proof-check validates checked-in packets for completeness and
local handoff metadata.
Repository layout
| Path | What it is |
| --- | --- |
| plugins/profile-and-optimize/skills/ | The Agent Skills, one directory per installed skill |
| plugins/profile-and-optimize/templates/skill/ | Starting point for a new skill |
| plugins/profile-and-optimize/server/ | MCP server, tool libraries, contract docs, and report renderer |
| plugins/profile-and-optimize/hooks/ | Runtime-neutral provenance guard plus the Claude Code adapter |
| configs/sol-ceilings.yaml | Datasheet-sourced hardware ceilings used by roofline reports |
| examples/workload-proof-packet/ | Synthetic fixture for packet and handoff validation |
| schemas/workload-proof-packet-v1.json | Public JSON Schema for workload proof packets |
| docs/METHODOLOGY.md | Measurement and reporting rules shared by the skills |
| mcp-descriptors/ | Offline MCP tool-schema snapshots used by skill lint |
Methodology
The skills enforce DRAFT and VERDICT labels, full performance context, asset
validation, and a clear next action. Read
docs/METHODOLOGY.md for the shared rules. The
performance hints adaptation
adds estimate-first triage, with a
synthetic example.
Optional integrations
A workflow system can consume workflow_handoff blocks when GPU workload
evidence needs to attach to a broader customer record. ProofPlane is one
possible consumer. It does not change this project's local contract,
validation gates, or runtime dependencies.
Development
- Add a skill from
plugins/profile-and-optimize/templates/skill/SKILL.md. - Add an MCP verb under
plugins/profile-and-optimize/server/and updatemcp_surface.py. - Read
CONTRIBUTING.mdbefore opening a pull request. - Run
make helpfor the local command reference.
Limitations
The project helps agents measure and report. It does not tune a cluster by itself. Every number depends on the workload and runtime context. Datasheet speed-of-light ceilings are upper bounds, not promises. External systems such as Grafana, Prometheus, GitHub, and zymtrace still need operator credentials and their own client configuration.
License
Project-authored material is MIT licensed. Third-party adaptations
retain their own terms and attribution in
THIRD_PARTY_NOTICES.md and LICENSES/.
Related Skills
claude-mem
97.5kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
93.0kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Understand-Anything
85.5kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.6kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
