nemo-mbridge-perf-activation-recompute
Validate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute.
Install / Use
npx skills add NVIDIA/skills --skill nemo-mbridge-perf-activation-recomputeInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Tags
Our assessment of nemo-mbridge-perf-activation-recompute
nemo-mbridge-perf-activation-recompute scores 93/100 on our quality scale, 533rd of 3,356 Development & Engineering skills we index (top 16%).
Its SKILL.md is 18 KB long, well organised into 21 sections with 2 code examples: a thorough specification that gives an agent plenty to work with.
With 3,421 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 5 days ago, so nemo-mbridge-perf-activation-recompute is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
nemo-mbridge-perf-activation-recompute compared with similar skills
All 4 of these similar skills score higher than nemo-mbridge-perf-activation-recompute; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| nemo-mbridge-perf-activation-recompute (this skill)by NVIDIA | 93 | 3.4k | 5d ago | SKILL.md |
| ai-job-searchby MadsLorentzen | 100 | 44.4k | today | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.7k | 2d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 6d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 6d ago | SKILL.md |
Frequently asked questions
- How do I install nemo-mbridge-perf-activation-recompute?
- Run
npx skills add NVIDIA/skills --skill nemo-mbridge-perf-activation-recompute. The install tabs above show the steps for each supported agent. - Which AI agents does nemo-mbridge-perf-activation-recompute work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is nemo-mbridge-perf-activation-recompute safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is nemo-mbridge-perf-activation-recompute still maintained?
- The repository was last updated 5 days ago, so nemo-mbridge-perf-activation-recompute is actively maintained.
Skill content
View source on GitHubname: nemo-mbridge-perf-activation-recompute description: >- Validate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute. Use for activation memory OOMs or regressions involving recompute_granularity, recompute_num_layers, recompute_modules, recompute_method, selective recompute, full recompute, or activation checkpointing. license: Apache-2.0
Activation Recompute
Stable docs: @docs/training/activation-recomputation.md Card: @skills/nemo-mbridge-perf-activation-recompute/card.yaml
<!-- Guidance refreshed: 2026-08-12. -->Activation recompute (activation checkpointing) trades additional forward work during backward for lower retained-activation memory. The useful checkpoint boundary depends on the model architecture, attention backend, parallelism, and the tensor that actually drives the per-rank peak.
Quick Decision Guide
- Confirm the pressure is real allocation, not allocator fragmentation. Compare
max_memory_allocated()withmax_memory_reserved()on every rank. - Keep an explicit no-recompute control when the workload fits. Under selective granularity,
recompute_modules=[]is valid and useful for this comparison. - Select the first boundary from the architecture and observed peak:
- Standard attention:
core_attnis the common first candidate. It is strongest when unfused attention materializes score/probability tensors. With Transformer Engine fused or Flash Attention, compare it against[]because those backends already rematerialize attention internals. - Multi-Latent Attention (MLA): start with
mla_up_projwhen expanded Q/K/V projections dominate. Addcore_attnonly when the attention-core state still matters. - Grouped MoE: start with
moe_actwhen the expert intermediate activation dominates; addlayernormwhen norm outputs are material. Use wholemoerecompute only after accounting for the extra expert compute and communication it replays. - Dense FFN:
mlpcan save the whole dense-MLP activation region, but it usually costs more compute than a narrow output-discard boundary.
- Standard attention:
- Change one label at a time. Record per-rank allocated/reserved peaks plus steady-state step time or throughput; do not infer a global module ranking from one recipe.
- Use full-layer recompute only when targeted selective boundaries do not make the workload fit. Full recompute has the broadest memory effect and the largest replay cost.
- Treat CUDA graphs, FP8, context-parallel communication, and overlap features as compatibility constraints, not afterthoughts.
Megatron Core's cpu_offloading=True is an alternative when PCIe/NVLink transfer overhead is preferable to replayed compute. It cannot be combined with activation recompute and is not compatible with pipeline parallelism greater than one.
Enablement
Selective recompute
cfg.model.recompute_granularity = "selective"
cfg.model.recompute_modules = ["core_attn"] # Common standard-attention candidate, not a universal default.
Use the decision table below to replace or extend that list for MLA, MoE, dense-MLP, or GDN workloads.
Full-layer recompute
cfg.model.recompute_granularity = "full"
cfg.model.recompute_method = "uniform"
cfg.model.recompute_num_layers = 1
uniform: checkpoint fixed groups ofrecompute_num_layerstransformer layers.block: checkpoint the firstrecompute_num_layerslayers on each pipeline stage, with virtual-pipeline-aware distribution.
Selective Module Decision Table
The currently pinned Megatron Core accepts these labels. A development branch can add model-specific labels, so validate against the exact target revision rather than copying a list across branches.
| Module | Checkpoint boundary | When to test it | Main cost or caveat |
|---|---|---|---|
| core_attn | Core attention | Standard attention, especially an unfused backend retaining attention intermediates | Replays attention. Incremental savings can be small with TE fused/Flash Attention; context parallelism can replay attention communication. |
| mla_up_proj | MLA Q/KV up-projection plus RoPE region | MLA models retaining expanded Q/K/V tensors | Replays the MLA expansion path. It is a distinct, potentially additive boundary from core_attn. |
| layernorm | Input and pre-MLP normalization outputs | Norm outputs contribute materially to the peak, often alongside MoE or MLA boundaries | Usually narrow, but savings depend on hidden size, sequence length, and which graph paths are active. |
| moe_act | Activation output between grouped expert FC1 and FC2 | Grouped MoE expert-intermediate activations dominate | Narrow output-discard checkpoint. It does not replay dispatch, FC1, or FC2, but has FP8 delayed-scaling restrictions. |
| mlp | Whole dense MLP | Dense layers dominate after narrower boundaries are exhausted | Replays the complete dense MLP. It has no effect on layers whose MLP is MoE. |
| moe | Whole MoE forward | A broad MoE region must be discarded to make the workload fit | Replays routing, dispatch/combine communication, experts, and shared-expert work. It is incompatible with expert-parallel overlap. |
| shared_experts | Non-overlapped shared-expert MLP | Shared experts are a distinct material peak | Replays the shared-expert MLP and is invalid with shared-expert overlap. Outer moe already removes its original-forward saves, but nesting can still change the transient backward-replay peak. |
| gdn_norm_out | GDN gated-normalization output | GDN/hybrid models retain this output | Replays the normalization and its HP-to-CP all-to-all path. |
For example, DeepSeek V4 configurations can use the model-specific mhc
label only with their required Megatron Core development branch. It is not a
portable label for the pinned revision and therefore is not included in the
table above.
Common performance configurations consequently fall into several patterns rather than one universal list:
- standard transformer recipes often use
core_attn; - MLA recipes often use
mla_up_proj, sometimes withmlp; - grouped-MoE recipes often use
moe_actorlayernormplusmoe_act; - higher-pressure MoE recipes sometimes use broader combinations such as
moepluslayernorm.
These are candidate patterns, not an ordering guarantee. Peak attribution and matched measurements decide the final list.
Measurement Contract
For every candidate, capture:
- exact Bridge and Megatron Core revisions;
- model, sequence length, micro/global batch sizes, precision, attention backend, and parallelism;
- the exact
recompute_granularity, module list, method, and layer count; - per-rank
max_memory_allocated()andmax_memory_reserved(); - steady-state step time or throughput after warmup;
- a short convergence or numerical-sanity check appropriate to the task.
Use a matched no-recompute control and change one recompute choice at a time. Peak memory from different jobs, backends, or parallel layouts is not a module-ranking benchmark.
Do not call a candidate successful merely because it advances farther than the control. Run through optimizer-state initialization and multiple steady-state steps: selective recompute can move the memory wall from forward into gradient synchronization or the optimizer without making the workload viable.
Matched H100 Evidence: Moonlight 16B
A 2026-08-12 short-run study used the exact Bridge revision
600d069b824dd5ce50367a311a5a3244478faf22 and Megatron Core revision
24bad8e677d22625d86ef2a54c9506b6e4992c93. The Moonlight 16B BF16
pretraining recipe ran on 8 H100 80GB GPUs with sequence length 4096, MBS=1,
GBS=4, TP=2, PP=1, CP=1, EP=8, mock data, and 20 steps. This model mixes one
dense layer with 26 MLA+MoE layers. Each row changed only
recompute_modules; all 20 losses were finite with zero skipped or NaN
iterations.
Peak allocated memory is the maximum post-optimizer value reported after iteration 2. Time and throughput are means over iterations 11--20.
| Selective modules | Peak allocated (GB) | Step time (ms) | TFLOP/s/GPU | Allocated vs [] | Time vs [] |
|---|---:|---:|---:|---:|---:|
| [] | 36.618 | 457.18 | 77.50 | control | control |
| core_attn | 36.614 | 474.80 | 74.44 | -0.01% | +3.85% |
| mla_up_proj | 35.902 | 480.89 | 73.72 | -1.96% | +5.19% |
| mla_up_proj, mlp | 35.917 | 496.73 | 71.50 | -1.91% | +8.65% |
| moe_act | 35.941 | 466.26 | 75.50 | -1.85% | +1.99% |
| layernorm, moe_act | 35.949 | 506.53 | 70.27 | -1.83% | +10.79% |
For this exact workload, moe_act is the best first boundary: it recovered
nearly as much allocated memory as mla_up_proj for less replay cost.
mla_up_proj is the next candidate if its roughly 39 MB additional reduction
matters. Adding mlp to mla_up_proj or layernorm to moe_act did not
improve the observed peak and made steps slower. Explicit core_attn added
cost without material memory benefit under fused attention.
Maximum reserved memory stayed near 40 GB and did not fall monotonically. That is allocator caching, not contrary evidence: boundary selection in this study is based on allocated memory and successful end-to-end steps.
Matched H100 Evidence: Nemotron 3 Nano
The same 2026-08-12 study used the native 16-H100 BF16 performance recipe for
the 52-layer hybrid Mamba/fused-attention MoE model. The matched short-run
configuration used sequence length 8192, MBS=1, GBS=16, TP=1, PP=1, CP=1,
EP=8, DP=16, expert-DP=2, HybridEP, grouped GEMM, TE CUDA graphs for attention and Mamba,
mock data, and 12 steps. Each row changed only recompute_modules.
| Selective modules | Outcome | Rank-0 measured peak | Failure or steady-state evidence |
|---|---|---:|---|
| [] | OOM after iteration 1 | 66.297 GB after iteration 1 | Iteration-2 MoE router allocation failed; hot ranks had about 72.9 GiB allocated. |
| core_attn | OOM in iteration 1 | not comparable | Grouped-expert linear allocation failed; explicit attention recompute did not make the fused-attention workload fit. |
| moe_act | OOM after iteration 1 | 62.103 GB after iteration 1 | 4.194 GB (6.33%) below the control at the matched checkpoint, but the iteration-2 output projection still needed 2 GiB. |
| layernorm, moe_act | OOM in iteration 1 | not comparable | Output projection still needed 2 GiB; CUDA-graph private pools were material. |
| moe | completed 12 steps | 64.653 GB after iteration 2 | 657.42 ms and 277.72 TFLOP/s/GPU over iterations 7--12. |
| moe, layernorm | completed 12 steps | 63.639 GB after iteration 2 | 677.62 ms and 270.62 TFLOP/s/GPU over iterations 7--12. |
Both successful rows had finite losses and zero skipped or NaN iterations.
For this exact capacity-limited recipe, whole-moe recompute is the smallest
tested passing boundary. Adding layernorm recovered another 1.014 GB (1.57%)
of rank-0 peak at 3.07% higher step time, so the recipe's broader combination
is justified when that headroom is required. Narrow moe_act produced real
activation relief but did not make the whole training step viable.
An exploratory native 8-H100 layout failed during FP32 optimizer-state initialization even at sequence length 4096. That is optimizer capacity, not a selective-boundary throughput baseline; no timing comparison from those runs is used here.
Cross-model conclusion
These measurements do not define one ranking. Moonlight fit with an empty
control and favored narrow moe_act; Nemotron required broad whole-moe
recompute; historical dense Llama evidence found whole-mlp replay costly and
lacked an empty control. The correct first candidate is therefore the narrowest
boundary implicated by the architecture and peak, followed by broader replay
only when the narrow choice does not pass the complete step.
Compatibility and Validation
Configuration semantics
recompute_granularity="selective"usesrecompute_modules;
Truncated for display — read the full file on GitHub.
Related Skills
ai-job-search
44.4kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.7kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
