NSDI 26 - SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling @UsenixOrg
NSDI 26 - SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling

Athinagoras Skiadopoulos, Stanford University; Mark Zhao, University of Colorado Boulder; Swapnil Gandhi, Stanford University and NVIDIA; Thomas Norrie, OpenAI; Shrijeet Mukherjee, NVIDIA; Christos Kozyrakis, Stanford University and NVIDIA

Mixture-of-Experts (MoE) models have become a widely-adopted solution to continue scaling model sizes without a corresponding linear increase in compute. During MoE model training, each input token is dynamically routed to a subset of experts—sparsely-activated feed-forward networks—within each transformer layer. The distribution of tokens assigned to each expert varies widely and rapidly over the course of training. To handle the wide load imbalance across experts, current systems are forced to either drop tokens assigned to popular experts, degrading convergence, or frequently rebalance resources allocated to each expert based on popularity, incurring high state migration overheads.
To break this performance-accuracy tradeoff, we introduce SYMI, an adaptive MoE training system. The key insight of SYMI is to decouple the placement of expert parameters from their large optimizer state. SYMI statically partitions the optimizer of each expert across all training nodes. Meanwhile, SYMI dynamically adjusts the placement of expert parameters by repurposing existing weight updates, avoiding migration overheads. In doing so, SYMI right-sizes the GPU resources allocated to each expert, on a per-iteration basis, with minimal overhead. Compared to state-of-the-art MoE training systems, DeepSpeed and FlexMoE, SYMI is able to achieve a 30.5% and 25.9% faster time-to-convergence, respectively.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State DecouplingPEPR 26 - Vision: Human-as-the-Unit Privacy Management with AI AgentsPEPR 26 - Surfacing Hidden Privacy Risks in Code: Lessons from LLM and Retrieval Assisted DetectionNSDI 26 - The GOODPUT System: A Machine Learning-Driven Optimization Framework for Dynamic...NSDI 26 - PrvTel: Lightweight Models for Private and Accurate Telemetry Data RetentionNSDI 26 - Count-Based Abstractions for Performance Verification of Contention PointsNSDI 26 - A Systematic Threat Analysis and Practical Attacks on Automated Frequency CoordinationNSDI 26 - ForestColl: Throughput-Optimal Collective Communications on Heterogeneous Network FabricsNSDI 26 - CacheCatalyst: Enhancing Web Caching for the Latency-Constrained InternetNSDI 26 - BURST: Seeking High-performance, Interoperability and Scalability in Soft-RDMANSDI 26 - MirrorNet: High-fidelity and Scalable Network Emulation for Software-defined WANPEPR 26 - Production Multi-Party Computation via the Distributed Aggregation Protocol
USENIX |

NSDI '26 - SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER