Source profileQuality 90/100

NVIDIA/skills/skills/nemo-mbridge-perf-activation-recompute/SKILL.md

nemo-mbridge-perf-activation-recompute

Validate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute. Use for activation memory OOMs or regressions involving recompute_granularity, recompute_num_layers, recompute_modules, recompute_method, selective recompute, full recompute, or activation checkpointing.

Source repository stars
3,106
Declared platforms
0
Static risk flags
0
Last source update
2026-08-25
Source checked
2026-08-26

Decision brief

What it does: where it fits

Stable docs: @docs/training/activation-recomputation.md Card: @skills/nemo-mbridge-perf-activation-recompute/card.yaml

Best for

    Not for

    • A module list is not portable across model families, attention backends, parallel layouts, or Megatron Core revisions.
    • Memory savings are nonlinear when boundaries overlap or nest; additive arithmetic is unreliable.

    Compatibility matrix

    Platform support, with evidence labels

    PlatformStatusEvidenceWhat to check
    CodexNot declaredNo explicit evidencePortability before use
    Claude CodeNot declaredNo explicit evidencePortability before use
    CursorNot declaredNo explicit evidencePortability before use
    Gemini CLINot declaredNo explicit evidencePortability before use
    Open the compatibility checker

    Installation

    Inspect first. Install second.

    The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

    Source-detected install commandSource
    npx skills add https://github.com/NVIDIA/skills --skill "skills/nemo-mbridge-perf-activation-recompute"
    Safe inspection promptEditorial

    Inspect the Agent Skill "nemo-mbridge-perf-activation-recompute" from https://github.com/NVIDIA/skills/blob/994b87022af46deada9fdb79fc560a77aaf931ce/skills/nemo-mbridge-perf-activation-recompute/SKILL.md at commit 994b87022af46deada9fdb79fc560a77aaf931ce. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

    Workflow

    What the source asks the agent to do

    1. 01

      Quick Decision Guide

      1. Confirm the pressure is real allocation, not allocator fragmentation. Compare maxmemoryallocated() with maxmemoryreserved() on every rank. 2. Keep an explicit no-recompute control when the workload fits. Under selective granularity, recomputemodules=[] is valid and useful for…

      Confirm the pressure is real allocation, not allocator fragmentation. Compare maxmemoryallocated() with maxmemoryreserved() on every rank.Keep an explicit no-recompute control when the workload fits. Under selective granularity, recomputemodules=[] is valid and useful for this comparison.Select the first boundary from the architecture and observed peak:
    2. 02

      Enablement

      Use the decision table below to replace or extend that list for MLA, MoE, dense-MLP, or GDN workloads.

      uniform: checkpoint fixed groups of recomputenumlayers transformer layers.block: checkpoint the first recomputenumlayers layers on each pipeline stage, with virtual-pipeline-aware distribution.Use the decision table below to replace or extend that list for MLA, MoE, dense-MLP, or GDN workloads.
    3. 03

      Selective recompute

      Use the decision table below to replace or extend that list for MLA, MoE, dense-MLP, or GDN workloads.

      Use the decision table below to replace or extend that list for MLA, MoE, dense-MLP, or GDN workloads.
    4. 04

      Full-layer recompute

      uniform: checkpoint fixed groups of recomputenumlayers transformer layers.

      uniform: checkpoint fixed groups of recomputenumlayers transformer layers.block: checkpoint the first recomputenumlayers layers on each pipeline stage, with virtual-pipeline-aware distribution.- uniform: checkpoint fixed groups of recomputenumlayers transformer layers. - block: checkpoint the first recomputenumlayers layers on each pipeline stage, with virtual-pipeline-aware distribution.
    5. 05

      Selective Module Decision Table

      The currently pinned Megatron Core accepts these labels. A development branch can add model-specific labels, so validate against the exact target revision rather than copying a list across branches.

      standard transformer recipes often use coreattn;MLA recipes often use mlaupproj, sometimes with mlp;grouped-MoE recipes often use moeact or layernorm plus moeact;

    Permission review

    Static risk signals and limitations

    No configured static risk pattern was detected

    This is not proof of safety. Runtime behavior, indirect dependencies, and hidden external systems are outside the static scan.

    Evidence record

    Why each signal appears

    EvidenceSourceComputedTestedEditorial
    SignalValueEvidence typeMeaning
    Quality score90/100ComputedDocumentation, specificity, maintenance, and trust rules
    Repository stars3,106SourceRepository attention, not individual Skill quality
    Compatibility0 platformsSourceDeclared in the catalog source record
    Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

    Pinned source

    Provenance and original SKILL.md

    Repository
    NVIDIA/skills
    Skill path
    skills/nemo-mbridge-perf-activation-recompute/SKILL.md
    Commit
    994b87022af46deada9fdb79fc560a77aaf931ce
    License
    Apache-2.0
    Collected
    2026-08-26
    Default branch
    main
    View the original SKILL.md

    Activation Recompute

    Stable docs: @docs/training/activation-recomputation.md Card: @skills/nemo-mbridge-perf-activation-recompute/card.yaml

    Activation recompute (activation checkpointing) trades additional forward work during backward for lower retained-activation memory. The useful checkpoint boundary depends on the model architecture, attention backend, parallelism, and the tensor that actually drives the per-rank peak.

    Quick Decision Guide

    1. Confirm the pressure is real allocation, not allocator fragmentation. Compare max_memory_allocated() with max_memory_reserved() on every rank.
    2. Keep an explicit no-recompute control when the workload fits. Under selective granularity, recompute_modules=[] is valid and useful for this comparison.
    3. Select the first boundary from the architecture and observed peak:
      • Standard attention: core_attn is the common first candidate. It is strongest when unfused attention materializes score/probability tensors. With Transformer Engine fused or Flash Attention, compare it against [] because those backends already rematerialize attention internals.
      • Multi-Latent Attention (MLA): start with mla_up_proj when expanded Q/K/V projections dominate. Add core_attn only when the attention-core state still matters.
      • Grouped MoE: start with moe_act when the expert intermediate activation dominates; add layernorm when norm outputs are material. Use whole moe recompute only after accounting for the extra expert compute and communication it replays.
      • Dense FFN: mlp can save the whole dense-MLP activation region, but it usually costs more compute than a narrow output-discard boundary.
    4. Change one label at a time. Record per-rank allocated/reserved peaks plus steady-state step time or throughput; do not infer a global module ranking from one recipe.
    5. Use full-layer recompute only when targeted selective boundaries do not make the workload fit. Full recompute has the broadest memory effect and the largest replay cost.
    6. Treat CUDA graphs, FP8, context-parallel communication, and overlap features as compatibility constraints, not afterthoughts.

    Megatron Core's cpu_offloading=True is an alternative when PCIe/NVLink transfer overhead is preferable to replayed compute. It cannot be combined with activation recompute and is not compatible with pipeline parallelism greater than one.

    Enablement

    Selective recompute

    cfg.model.recompute_granularity = "selective"
    cfg.model.recompute_modules = ["core_attn"]  # Common standard-attention candidate, not a universal default.
    

    Use the decision table below to replace or extend that list for MLA, MoE, dense-MLP, or GDN workloads.

    Full-layer recompute

    cfg.model.recompute_granularity = "full"
    cfg.model.recompute_method = "uniform"
    cfg.model.recompute_num_layers = 1
    
    • uniform: checkpoint fixed groups of recompute_num_layers transformer layers.
    • block: checkpoint the first recompute_num_layers layers on each pipeline stage, with virtual-pipeline-aware distribution.

    Selective Module Decision Table

    The currently pinned Megatron Core accepts these labels. A development branch can add model-specific labels, so validate against the exact target revision rather than copying a list across branches.

    ModuleCheckpoint boundaryWhen to test itMain cost or caveat
    core_attnCore attentionStandard attention, especially an unfused backend retaining attention intermediatesReplays attention. Incremental savings can be small with TE fused/Flash Attention; context parallelism can replay attention communication.
    mla_up_projMLA Q/KV up-projection plus RoPE regionMLA models retaining expanded Q/K/V tensorsReplays the MLA expansion path. It is a distinct, potentially additive boundary from core_attn.
    layernormInput and pre-MLP normalization outputsNorm outputs contribute materially to the peak, often alongside MoE or MLA boundariesUsually narrow, but savings depend on hidden size, sequence length, and which graph paths are active.
    moe_actActivation output between grouped expert FC1 and FC2Grouped MoE expert-intermediate activations dominateNarrow output-discard checkpoint. It does not replay dispatch, FC1, or FC2, but has FP8 delayed-scaling restrictions.
    mlpWhole dense MLPDense layers dominate after narrower boundaries are exhaustedReplays the complete dense MLP. It has no effect on layers whose MLP is MoE.
    moeWhole MoE forwardA broad MoE region must be discarded to make the workload fitReplays routing, dispatch/combine communication, experts, and shared-expert work. It is incompatible with expert-parallel overlap.
    shared_expertsNon-overlapped shared-expert MLPShared experts are a distinct material peakReplays the shared-expert MLP and is invalid with shared-expert overlap. Outer moe already removes its original-forward saves, but nesting can still change the transient backward-replay peak.
    gdn_norm_outGDN gated-normalization outputGDN/hybrid models retain this outputReplays the normalization and its HP-to-CP all-to-all path.

    For example, DeepSeek V4 configurations can use the model-specific mhc label only with their required Megatron Core development branch. It is not a portable label for the pinned revision and therefore is not included in the table above.

    Common performance configurations consequently fall into several patterns rather than one universal list:

    • standard transformer recipes often use core_attn;
    • MLA recipes often use mla_up_proj, sometimes with mlp;
    • grouped-MoE recipes often use moe_act or layernorm plus moe_act;
    • higher-pressure MoE recipes sometimes use broader combinations such as moe plus layernorm.

    These are candidate patterns, not an ordering guarantee. Peak attribution and matched measurements decide the final list.

    Measurement Contract

    For every candidate, capture:

    • exact Bridge and Megatron Core revisions;
    • model, sequence length, micro/global batch sizes, precision, attention backend, and parallelism;
    • the exact recompute_granularity, module list, method, and layer count;
    • per-rank max_memory_allocated() and max_memory_reserved();
    • steady-state step time or throughput after warmup;
    • a short convergence or numerical-sanity check appropriate to the task.

    Use a matched no-recompute control and change one recompute choice at a time. Peak memory from different jobs, backends, or parallel layouts is not a module-ranking benchmark.

    Do not call a candidate successful merely because it advances farther than the control. Run through optimizer-state initialization and multiple steady-state steps: selective recompute can move the memory wall from forward into gradient synchronization or the optimizer without making the workload viable.

    Matched H100 Evidence: Moonlight 16B

    A 2026-08-12 short-run study used the exact Bridge revision 600d069b824dd5ce50367a311a5a3244478faf22 and Megatron Core revision 24bad8e677d22625d86ef2a54c9506b6e4992c93. The Moonlight 16B BF16 pretraining recipe ran on 8 H100 80GB GPUs with sequence length 4096, MBS=1, GBS=4, TP=2, PP=1, CP=1, EP=8, mock data, and 20 steps. This model mixes one dense layer with 26 MLA+MoE layers. Each row changed only recompute_modules; all 20 losses were finite with zero skipped or NaN iterations.

    Peak allocated memory is the maximum post-optimizer value reported after iteration 2. Time and throughput are means over iterations 11--20.

    Selective modulesPeak allocated (GB)Step time (ms)TFLOP/s/GPUAllocated vs []Time vs []
    []36.618457.1877.50controlcontrol
    core_attn36.614474.8074.44-0.01%+3.85%
    mla_up_proj35.902480.8973.72-1.96%+5.19%
    mla_up_proj, mlp35.917496.7371.50-1.91%+8.65%
    moe_act35.941466.2675.50-1.85%+1.99%
    layernorm, moe_act35.949506.5370.27-1.83%+10.79%

    For this exact workload, moe_act is the best first boundary: it recovered nearly as much allocated memory as mla_up_proj for less replay cost. mla_up_proj is the next candidate if its roughly 39 MB additional reduction matters. Adding mlp to mla_up_proj or layernorm to moe_act did not improve the observed peak and made steps slower. Explicit core_attn added cost without material memory benefit under fused attention.

    Maximum reserved memory stayed near 40 GB and did not fall monotonically. That is allocator caching, not contrary evidence: boundary selection in this study is based on allocated memory and successful end-to-end steps.

    Matched H100 Evidence: Nemotron 3 Nano

    The same 2026-08-12 study used the native 16-H100 BF16 performance recipe for the 52-layer hybrid Mamba/fused-attention MoE model. The matched short-run configuration used sequence length 8192, MBS=1, GBS=16, TP=1, PP=1, CP=1, EP=8, DP=16, expert-DP=2, HybridEP, grouped GEMM, TE CUDA graphs for attention and Mamba, mock data, and 12 steps. Each row changed only recompute_modules.

    Selective modulesOutcomeRank-0 measured peakFailure or steady-state evidence
    []OOM after iteration 166.297 GB after iteration 1Iteration-2 MoE router allocation failed; hot ranks had about 72.9 GiB allocated.
    core_attnOOM in iteration 1not comparableGrouped-expert linear allocation failed; explicit attention recompute did not make the fused-attention workload fit.
    moe_actOOM after iteration 162.103 GB after iteration 14.194 GB (6.33%) below the control at the matched checkpoint, but the iteration-2 output projection still needed 2 GiB.
    layernorm, moe_actOOM in iteration 1not comparableOutput projection still needed 2 GiB; CUDA-graph private pools were material.
    moecompleted 12 steps64.653 GB after iteration 2657.42 ms and 277.72 TFLOP/s/GPU over iterations 7--12.
    moe, layernormcompleted 12 steps63.639 GB after iteration 2677.62 ms and 270.62 TFLOP/s/GPU over iterations 7--12.

    Both successful rows had finite losses and zero skipped or NaN iterations. For this exact capacity-limited recipe, whole-moe recompute is the smallest tested passing boundary. Adding layernorm recovered another 1.014 GB (1.57%) of rank-0 peak at 3.07% higher step time, so the recipe's broader combination is justified when that headroom is required. Narrow moe_act produced real activation relief but did not make the whole training step viable.

    An exploratory native 8-H100 layout failed during FP32 optimizer-state initialization even at sequence length 4096. That is optimizer capacity, not a selective-boundary throughput baseline; no timing comparison from those runs is used here.

    Cross-model conclusion

    These measurements do not define one ranking. Moonlight fit with an empty control and favored narrow moe_act; Nemotron required broad whole-moe recompute; historical dense Llama evidence found whole-mlp replay costly and lacked an empty control. The correct first candidate is therefore the narrowest boundary implicated by the architecture and peak, followed by broader replay only when the narrow choice does not pass the complete step.

    Compatibility and Validation

    Configuration semantics

    • recompute_granularity="selective" uses recompute_modules; an empty list is accepted as an explicit control.
    • recompute_granularity="full" uses recompute_method and recompute_num_layers; selective labels do not apply.
    • Full granularity supersedes selective module choices rather than composing with them.
    • Unknown labels fail Megatron Core validation. Labels may differ on development branches, so use the exact revision's TransformerConfig validator as the source of truth.

    Attention backend and context parallelism

    • TE fused and Flash Attention already use internal rematerialization. Explicit core_attn may still change retained inputs/outputs, but it must earn its place in a matched [] comparison.
    • Under context parallelism, an attention checkpoint can replay communication as well as compute. Include CP size and topology in the measurement record.

    MoE restrictions

    • Whole-moe recompute is incompatible with expert-parallel overlap because backward replay would repeat the overlapped routing/communication region.
    • shared_experts recompute is incompatible with shared-expert overlap.
    • moe_act applies to grouped-GEMM experts and is the narrower choice when only the expert activation needs to be discarded.
    • mlp targets dense MLPs and is a no-op on MoE layers; mixed dense/MoE models can still benefit on their dense layers.

    FP8 restrictions

    • moe_act and layernorm recompute are not supported with FP8 delayed scaling and require a compatible Transformer Engine version.
    • Absorbed MLA paths have additional FP8/FP4 restrictions. Validate the exact model/provider path before selecting mla_up_proj.

    CUDA graphs

    • Selective recompute is valid only when a checkpointed module lies wholly inside or wholly outside the selected graph scope. A checkpoint boundary that straddles a graph boundary is invalid.
    • Capture/warmup can bypass checkpoint wrappers, so verify the final graph scope and replay path rather than assuming eager behavior carries over.
    • Full recompute with CUDA graphs requires cuda_graph_impl="full_iteration" in the pinned Megatron Core. Otherwise disable CUDA graphs; scoped/local graph capture is not a substitute for full-iteration capture here.

    Historical Measurement: Context, Not a Module Ranking

    Historical H100 measurements from Bridge PR #3107 used Llama 3 70B SFT on 32 H100 80GB GPUs with FP8 current scaling, sequence length 4096, micro-batch size 1, global batch size 32, TP=4, PP=4, VPP=5, and DP=2:

    ConfigurationTFLOP/s/GPUPeak memory
    core_attn baseline in that run~70458.8 GB (OOM on rank 0)
    mlp593.655.6 GB
    mlp + core_attn586.855.6 GB
    core_attn + layernorm~70259.6 GB (OOM on rank 0)
    Golden throughput recorded in the PR context709.93Not a paired memory measurement

    Limitations of this evidence:

    • it did not include a matched no-recompute row;
    • the golden row was not a paired module-only comparison;
    • the measurements cover one dense Llama workload, not MLA or MoE;
    • the table supports the local memory/throughput tradeoff only and must not be used to rank all recompute labels.

    Code Anchors

    • Selective-label validation and cross-feature checks: 3rdparty/Megatron-LM/megatron/core/transformer/transformer_config.py
    • Checkpoint implementations: 3rdparty/Megatron-LM/megatron/core/tensor_parallel/random.py
    • Standard-attention checkpoint boundary: 3rdparty/Megatron-LM/megatron/core/transformer/attention.py
    • MLA up-projection boundary: 3rdparty/Megatron-LM/megatron/core/transformer/multi_latent_attention.py
    • Layernorm, dense-MLP, and outer-MoE placement: 3rdparty/Megatron-LM/megatron/core/transformer/transformer_layer.py
    • Grouped expert activation boundary: 3rdparty/Megatron-LM/megatron/core/transformer/moe/experts.py
    • Shared-expert and whole-MoE paths: 3rdparty/Megatron-LM/megatron/core/transformer/moe/moe_layer.py
    • GDN normalization boundary: 3rdparty/Megatron-LM/megatron/core/ssm/gated_delta_net/gdn.py

    Failure Diagnosis

    SymptomLikely causeNext action
    core_attn gives little or no peak reductionFused/Flash attention already rematerializes the expensive internals, or the peak is elsewhereCompare with [], attribute the peak, then test the architecture-specific boundary such as mla_up_proj or moe_act.
    MLA still OOMs after core_attnExpanded Q/K/V projection tensors, not attention-core tensors, dominateTest mla_up_proj; add core_attn only if matched evidence supports it.
    MoE peak remains highExpert intermediate or norm outputs dominateTest moe_act, then layernorm; reserve whole moe for broader pressure.
    Expert-overlap validation failsWhole-moe or shared_experts recompute conflicts with overlapKeep overlap and use a compatible inner boundary, or disable overlap and remeasure the entire configuration.
    A selected label has no measurable effectThat module is absent or inactive on the measured layers, or graph capture bypassed the wrapperInspect the provider/layer mix and final graph scope; for example, mlp is ineffective on pure-MoE layers.
    Full recompute plus CUDA graphs assertsGraph implementation is not full-iterationSet cuda_graph_impl="full_iteration" or disable CUDA graphs.
    Reserved memory is high but allocated memory is stableAllocator fragmentation or cachingTry PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True before adding recompute.
    OOM moves to a different rank after enabling recomputePipeline/virtual-pipeline layer distribution changed the bottleneckCompare per-rank peaks and tune full block/uniform placement or selective boundaries for the actual hot stage.
    A candidate gets farther but still OOMsRecompute moved the peak into gradient synchronization or optimizer-state initializationRecord the changed failure stage as diagnostic evidence, but require optimizer initialization and multiple steady steps before calling it a pass.

    Known Limitations

    • A module list is not portable across model families, attention backends, parallel layouts, or Megatron Core revisions.
    • Memory savings are nonlinear when boundaries overlap or nest; additive arithmetic is unreliable.
    • Full recompute changes RNG execution paths; dropout workloads need a numerical/convergence check.
    • Activation recompute does not address parameter, optimizer-state, or allocator-fragmentation pressure.
    • The correct result is the smallest measured replay cost that satisfies the per-rank memory target, not the longest module list.

    Further Reading

    Frequently asked questions

    What to verify before installation and use

    What does the nemo-mbridge-perf-activation-recompute source document cover?

    Stable docs: @docs/training/activation-recomputation.md Card: @skills/nemo-mbridge-perf-activation-recompute/card.yaml

    How do I install nemo-mbridge-perf-activation-recompute?

    The source record exposes this install command: npx skills add https://github.com/NVIDIA/skills --skill "skills/nemo-mbridge-perf-activation-recompute". Inspect the command and pinned source before running it.

    Alternatives

    Compare before choosing

    Computed 10045,643

    coreyhaines31/marketingskills

    ab-testing

    When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program

    Computed 10029,095

    garrytan/gbrain

    bulk-ingestion

    End-to-end discipline for turning any large data source (audio libraries, email takeouts, document corpora, chat exports, API dumps) into brain pages at scale. The lifecycle spine: SCHEMA → ACCESS → TRIAL → EVALUATE → IMPROVE → CODIFY → TEST → SKILLIFY → BULK → MONITOR. State is tracked in a durable JSON manifest (see MANIFEST-PATTERN.md) so any crash, session boundary, or subagent fan-out resumes from ground truth instead of memory.

    Computed 10024,975

    alirezarezvani/claude-skills

    app-store-optimization

    App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklist

    Computed 1005,248

    dotnet/skills

    migrate-vstest-to-mtp

    Migrates .NET test projects from VSTest to Microsoft.Testing.Platform (MTP). Use when user asks to "migrate to MTP", "switch from VSTest", "enable Microsoft.Testing.Platform", "use MTP runner", set OutputType=Exe only for test projects in Directory.Build.props, or mentions EnableMSTestRunner, EnableNUnitRunner, or UseMicrosoftTestingPlatformRunner. USE FOR: MTP behavioral differences vs VSTest (exit code 8, zero tests discovered, --ignore-exit-code, TESTINGPLATFORM_EXITCODE_IGNORE); centralizing