matlab-optimize-gpu-codegen
Optimize MATLAB design files for GPU Coder to generate faster CUDA code. Iteratively profiles, rewrites, and benchmarks until performance targets are met or diagnostics are resolved
Install / Use
npx skills add matlab/matlab-agentic-toolkit --skill matlab-optimize-gpu-codegenInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Our assessment of matlab-optimize-gpu-codegen
matlab-optimize-gpu-codegen scores 93/100 on our quality scale, 804th of 4,646 Development & Engineering skills we index (top 18%).
Its SKILL.md is 18 KB long, well organised into 20 sections with 10 code examples: a thorough specification that gives an agent plenty to work with.
With 1,098 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 18 days ago, so matlab-optimize-gpu-codegen is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
matlab-optimize-gpu-codegen compared with similar skills
All 4 of these similar skills score higher than matlab-optimize-gpu-codegen; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| matlab-optimize-gpu-codegen (this skill)by matlab | 93 | 1.1k | 18d ago | SKILL.md |
| ai-job-searchby MadsLorentzen | 100 | 44.9k | today | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.7k | 3d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
Frequently asked questions
- How do I install matlab-optimize-gpu-codegen?
- Run
npx skills add matlab/matlab-agentic-toolkit --skill matlab-optimize-gpu-codegen. The install tabs above show the steps for each supported agent. - Which AI agents does matlab-optimize-gpu-codegen work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is matlab-optimize-gpu-codegen safe to use?
- It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is matlab-optimize-gpu-codegen still maintained?
- The repository was last updated 18 days ago, so matlab-optimize-gpu-codegen is actively maintained.
Skill content
View source on GitHubname: matlab-optimize-gpu-codegen description: > Optimize MATLAB design files for GPU Coder to generate faster CUDA code. Iteratively profiles, rewrites, and benchmarks until performance targets are met or diagnostics are resolved. Use when asked to: optimize for GPU Coder, improve GPU codegen performance, profile generated GPU/CUDA code, profile GPU MEX, fix gpuPerformanceAnalyzer diagnostics, speed up GPU MEX, reduce GPU memory transfers, improve kernel parallelism, rewrite MATLAB for CUDA, or run gpuPerformanceAnalyzer. license: https://www.mathworks.com/content/dam/mathworks/license/pmrl/license.md metadata: author: MathWorks version: "1.1"
Optimize MATLAB for GPU Code Generation
Iteratively optimize a MATLAB design file for GPU Coder: compile, benchmark, apply structural optimizations, profile with gpuPerformanceAnalyzer, fix diagnostics, and verify numerical equivalence at every step.
When to Use
- User has a MATLAB function and wants faster GPU MEX or CUDA code
- User mentions GPU Coder, codegen, CUDA, gpuPerformanceAnalyzer
- User asks to profile generated GPU/CUDA code or GPU MEX (profiling generated GPU code is the entry point to this skill's diagnostic-fix workflow)
- User wants to reduce GPU memory, improve kernel parallelism, or fix Performance Analyzer diagnostics
- User has a
.mdesign file and representative inputs
When NOT to Use
- Workflows with no codegen
coder.gpuConfig("exe")— standalone executables cannot be benchmarked or equivalence-checked from MATLAB. Suggest the user switch tomexorlib/dlland regenerateexefrom the final optimized source.- Simulink GPU code generation
- Writing new MATLAB functions from scratch (this skill optimizes existing code)
- Optimizing helper functions called by the design file — this skill optimizes the main design file only
- Hardware setup or CUDA toolkit installation
- General MATLAB performance tuning without GPU involvement
- The user's prompt does not mention GPU, codegen, CUDA, MEX, or profiling. Activation must be driven by the user's prompt alone — do not infer GPU intent from filenames or function contents. If unsure, ask before activating.
Workflow
Setup — Create Session Directory
All codegen artifacts, profiling outputs, and optimized versions go in a
single temp directory for the entire session. pwd must be sessionDir for
every codegen/benchmark/PA call — otherwise compiled MEX and SIL binaries
land in the user's working directory.
Two helpers this skill calls — benchmarkMex and extractDiagnostics — live
in the scripts/ subfolder of this skill (the folder containing this
SKILL.md). A MATLAB function is only callable by name when its folder is on the
path, so add scripts/ to the path in Setup and remove it in the cleanup. Then
call the helpers by bare name (benchmarkMex(...), extractDiagnostics(...)).
% TEMPLATE — not executable
sessionDir = fullfile(tempdir, "gpu_opt_" + string(datetime("now", Format="yyyyMMdd_HHmmss")));
mkdir(sessionDir);
scriptsDir = fullfile("<skill_dir>", "scripts"); % this skill's scripts/ folder (absolute)
addedDir = fileparts(which("<designFile>"));
addpath(scriptsDir, addedDir);
oldDir = cd(sessionDir);
cleanupCd = onCleanup(@() (cd(oldDir), rmpath(scriptsDir), rmpath(addedDir))); %#ok<NASGU>
Write all artifacts to sessionDir only — never to the user's working directory.
Step 1 — Codegen on Original
Run codegen on the unmodified design file to discover what actually fails.
Do NOT guess which functions are unsupported — let the compiler tell you.
1a. Pick the codegen config. Use the config the user provides. If none specified, default to MEX:
cfg = coder.gpuConfig("mex"); % default — replace if user specifies a config
Supported targets: mex, lib, dll. The exe target is not
supported by this skill, so Steps 2–5 cannot benchmark or verify equivalence.
If the user provides coder.gpuConfig("exe"), stop and ask them to either:
- switch to
mexfor the optimization workflow (recommended — fastest iteration), or - switch to
lib/dllif they need a deployable artifact (the skill will enable SIL to benchmark via a generated MEX).
Once optimization is complete, the user can regenerate with exe from the
final optimized source.
1b. For lib/dll configs: enable SIL. This is mandatory — without SIL there is no callable MEX, so Steps 2–5 cannot benchmark. Set this before calling codegen:
% TEMPLATE — not executable
if ~isa(cfg, 'coder.MexCodeConfig')
cfg.VerificationMode = 'SIL'; % required for benchmarking lib/dll targets
end
If SIL fails or is unavailable (e.g., Embedded Coder license missing), report this and stop — do not silently skip benchmarking.
1c. Resolve inputs. If the user provided concrete input values, use them
as-is. If the user provided only types/sizes (e.g., "two double vectors of
size 1024x1"), synthesize inputs matching the spec — record exactly what
you generated (type, size, location, generator) so the Final Report can list
it. A reasonable default is randn with a fixed rng seed for floats,
randi for integers, rand > 0.5 for logicals; keep inputs on the CPU
unless the user said otherwise or PA later flags UseGpuInput. If the user
gave neither values nor types/sizes, ask for representative sizes and types —
codegen -args needs a concrete signature and the input shape drives which
optimizations win.
1d. Run codegen:
% TEMPLATE — not executable
codegen -config cfg <designFile> -args {<inputs>}
If codegen fails, read the errors and fix only what is reported as unsupported.
Save the fixed file as <designFile>_v1.m in sessionDir.
v1 must successfully codegen. Re-run codegen until it passes. This
produces the baseline MEX: <designFile>_v1_mex (for lib/dll configs, the
SIL-generated MEX has the same name and is callable identically).
Step 2 — Baseline Benchmark
Use the MEX generated in Step 1 directly — do not re-codegen. Benchmark with the convergence-based helper:
% TEMPLATE — not executable
baseline = benchmarkMex("<designFile>_v1_mex", {<inputs>});
baselineTime = baseline.MedianTime;
fprintf("Baseline: %.4f ms\n", baselineTime*1000);
Rules:
- Always use
benchmarkMex(orgputimeit) — never usetic/tocfor GPU timing (GPU ops are async) benchmarkMexhandles warmup and convergence automatically- Record
baselineTime— all improvements are measured against this
Step 3 — Structural Optimization Loop
Apply optimizations iteratively. Each iteration:
- Create
<designFile>_v<N>.minsessionDirwith the next optimization - Verify numerical equivalence against the original (multi-input — see Step 3b)
- Run codegen to
sessionDir— if it fails, fix or revert - Benchmark the new MEX with
benchmarkMex— compare against best so far - Keep the fastest passing version as the current best
MATLAB-level restructuring patterns (rewrites that improve codegen regardless of which GPU primitive you eventually choose):
| Pattern in Source | Optimization | Effect on Generated CUDA |
|---|---|---|
| Array-valued expression inside a loop that does not depend on the loop variable (e.g., w = weights / sum(weights) recomputed every iteration) | Hoist the expression above the loop | Eliminates redundant per-thread computation. GPU Coder cannot always prove loop-invariance for array expressions — it inlines them into the kernel, so every thread recomputes the same work. |
| Implicit expansion could replace an explicit loop (e.g., for i=1:N, out(i,:) = A(i,:) + B; end where B is a row vector) | Replace the loop with out = A + B using implicit expansion | GPU Coder generates a single fused kernel for implicit expansion. The explicit loop may also parallelize, but implicit expansion produces cleaner kernels with no loop overhead. |
| Divergent branching inside a parallelizable loop (e.g., if x(i) > 0, a = f(x(i)); else, a = g(x(i)); end where the condition varies unpredictably across elements) | Where both branches are cheap, compute both and select with a mask (e.g., a = mask.*f(x) + (~mask).*g(x)) | Divergent if/else causes warp divergence — threads in the same warp serialize across branches. A branchless mask keeps every thread on the same instruction path, restoring full warp throughput. |
The table above covers MATLAB-level restructuring only; it does not
enumerate GPU Coder primitives. Before committing to any optimization, read
references/gpu-codegen-functions.md end-to-end — it documents kernel
pragmas, parallel reductions and scans, atomics, stencils, and memory
placement. The right primitive depends on the loop's data-flow shape
(where the result lives, how dependencies chain, whether iterations
collide); the closest-looking table row is often not the right answer.
Match semantics to the loop, not surface appearance.
Exit criteria for this loop:
- Performance target met (if user specified one)
- No more structural optimizations the agent can identify
- Maximum 5 structural iterations (Step 5's diagnostic-fix loop has its own separate cap of 5 — the two are independent)
Step 3b — Verify Numerical Equivalence
Run both original and optimized on up to 5 input sets. Use the user's
original inputs as the baseline, then generate variants matching the same
types and sizes. Match the generator to the input type (randn for float,
randi for integer, rand > 0.5 for logical) and keep values in the type's
valid range.
% TEMPLATE — not executable — float case; adapt generator to the input type
% Input sets — variants are built from in1, the first input resolved in Step 1c
in1 = <user_or_synthesized_input_from_step_1c>; % original example
rng(42); in3 = randn(size(in1), 'like', in1); % random
rng(99); in4 = randn(size(in1), 'like', in1) * 1e6; % large magnitude
Include an edge-case input only where it is valid for this function — an all-zeros set exercises little and is degenerate if the function divides by something derived from the input. Pick a variant that meaningfully stresses the function instead.
Compare each output independently:
% TEMPLATE — not executable
outOrig = <designFile>(testInput);
outOpt = <designFile>_v<N>(testInput);
% Gather gpuArray outputs to CPU for comparison
if isa(outOpt, 'gpuArray'), outOpt = gather(outOpt); end
% If outputs can contain NaN, first check NaN positions match (see NaN rule
% below) — max() ignores NaN, so a corrupted NaN would otherwise pass silently.
denom = max(abs(outOrig(:)));
if denom == 0
relErr = max(abs(outOpt(:))); % absolute error when expected is zero
else
relErr = max(abs(outOrig(:) - outOpt(:))) / denom;
end
fprintf("relErr = %.2e\n", relErr);
Tolerance by output type:
| Output Type | Tolerance |
|-------------|-----------|
| double | 1e-6 |
| single | 1e-3 |
| half | 1e-2 |
| Integer / logical | Exact (isequal) |
Rules:
- Always compare against the original unmodified function, not the previous version
- If the function has multiple outputs, check each with its own type-appropriate tolerance
- If output types differ between original and optimized, that's a bug — fix it
- Normalize relative error by
max(abs(expected)), not element-wise - Handle NaN: check NaN positions match (
isequal(isnan(a), isnan(b))), then compare non-NaN elements - gpuArray wrapping of inputs is allowed (e.g., after UseGpuInput fix) — it doesn't change the function contract
Step 4 — Profile with gpuPerformanceAnalyzer
Always run this step unless the user specified a performance target and it was met in Step 3. PA reveals issues (memory copies, low parallelism) that benchmarks alone cannot detect.
Run PA on the current best version to find remaining bottlenecks. Pass the
same cfg from Step 1 so PA profiles the same target the user is
optimizing for:
Truncated for display — read the full file on GitHub.
Related Skills
ai-job-search
44.9kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.7kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
