nemo-mbridge-perf-memory-tuning
Covers expandable segments, PEFT and SP input re-gather, parallelism resizing, and activation recomputation. Includes CPU offloading constraints and common out-of-memory fixes.
Your agent can do more. Find the skill that makes it happen.
Find your next superpower If you are an agent, refer to our llms.txt for full access.Covers expandable segments, PEFT and SP input re-gather, parallelism resizing, and activation recomputation. Includes CPU offloading constraints and common out-of-memory fixes.
Covers dispatch and combine overlap for MoE expert-parallel communication. Includes flex dispatcher backends and expert weight-gradient scheduling.
Guides selection among alltoall, DeepEP, and HybridEP based on hardware, expert-parallel degree, and optimization stage. Summarizes patterns from DSV3, Qwen3, Qwen3-Next, and VLM bring-up work.
Provides representative, point-in-time MoE training configurations by hardware and model family. Calls for revalidating runtime, semantics, topology, and steady-state throughput before relying on them.
Covers context-parallel sizing, selective recomputation, and dispatcher choices. Draws practical patterns from DSV3, Qwen3, and Qwen3-Next long-context experiments.
Provides an evidence-gated optimization workflow using measurement contracts, the Three Walls framework, and parallel folding. Covers profiling, matched A/B tuning, and final validation.
Compares FSDP and 3D-parallel approaches to MoE VLM training. Uses lessons from Qwen3-VL, Qwen3-Next, and other multimodal experiments.
Provides sizing rules and hardware topology mapping for parallelism strategies. Covers configuring combined parallelism in Megatron Bridge.
Covers offline LLM packing, collate-time VLM packing, and Energon online packing. Includes validation and context-parallel constraints for packed sequences and long-context training.
Provides operational guidance for TP, DP, and PP communication overlap. Includes configuration controls, relevant code locations, pitfalls, and verification.
Selects and customizes library and benchmark recipes for GPU counts, hardware, sequence lengths, and pretraining, SFT, or PEFT goals. Covers parallelism resizing and distinguishes convergence changes, execution tuning, and benchmark-only shortcuts.
Covers fault tolerance, straggler detection, and in-process restart in Megatron Bridge. Includes preemption handling and the re-run state machine.
Skills give your agent reusable instructions for a specific job. Pick one, read what it does, and bring it into your workflow.
A name only tells half the story. Search the full description and instructions to find the right fit.
Read the skill, visit its source, and see exactly what you’re adding to your agent.
Copy the install command from a skill page and run it in your project.
npx skillycli add owner/repo --skill name