NSDI 26 - Sparse Checkpointing for Fast and Reliable MoE Training @UsenixOrg
NSDI 26 - Sparse Checkpointing for Fast and Reliable MoE Training  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
Sparse Checkpointing for Fast and Reliable MoE Training

Swapnil Gandhi, Stanford University; Christos Kozyrakis, Stanford University and NVIDIA

As large language models scale, training them requires thousands of GPUs over extended durations—making frequent failures an inevitable reality. While checkpointing remains the primary fault-tolerance mechanism, existing methods fall short when applied to Mixture-of-Experts (MoE) models. Due to their substantially larger training state, MoE models exacerbate checkpointing overheads, often causing costly stalls or prolonged recovery that severely degrade training efficiency.
We present MoEvement, a distributed, in-memory checkpointing system tailored for MoE models. MoEvement is built on three key ideas: (1) sparse checkpointing, which incrementally snapshots subsets of experts across iterations to reduce overhead; (2) a sparse-to-dense checkpoint conversion mechanism that incrementally reconstructs consistent dense checkpoints from sparse snapshots; and (3) upstream logging of activations and gradients at pipeline-stage boundaries, enabling localized recovery without re-executing unaffected workers. Evaluations across diverse MoE models with up to 64 experts show that MoEvement reduces checkpointing overhead by up to 4× and recovery overhead by up to 31× compared to state-of-the-art approaches, sustaining ETTR ≥ 0.94 even under frequent failures (MTBF as low as 10 minutes) and delivering up to 8× overall training speedup, all without compromising synchronous training semantics. Overall, MoEvement offers a scalable, practical fault-tolerance solution for the next generation of sparsely activated models.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - Sparse Checkpointing for Fast and Reliable MoE TrainingVehicleSec 25 - CarPlay at Risk: Unveiling Security Threats of Third-Party Infotainment AdaptersPEPR 26 - The Emperors New Embeddings: Obfuscating ML Inputs Doesnt Provide PrivacyNSDI 26 - RLBoost: Harvesting Preemptible Cloud Resources for Cost-Efficient Reinforcement LearningPEPR 26 - CA-CI: A Normative Framework for Evaluating Privacy and Dignity in AI GovernanceNSDI 26 - Decoding RSSI Compression in RFID: Dynamic RCS Modeling and Tag-Intrinsic Power Metrics..NSDI 26 - Over-Threshold Multiparty Private Set Intersection for Collaborative...SREcon24 Europe/Middle East/Africa - Dude, You Forgot the Feedback: How Your Open Loop Control...NSDI 26 - Queue-Mem: Energy-Efficient Hardware Storage for Advanced Network Function AccelerationPEPR 26 - Dismantling the Barriers to Personal Data PortabilityNSDI 26 - Defending against Traffic Analysis Attacks with Flexible In-Network ObfuscationUSENIX Security 24 - HYPERPILL: Fuzzing for Hypervisor-bugs by Leveraging the Hardware...
USENIX |

NSDI '26 - Sparse Checkpointing for Fast and Reliable MoE Training

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER