Uploaded June 2026 | Updated September 2026, 3 weeks ago
Yiran Lei, Carnegie Mellon University and MangoBoost; Dongjoo Lee, MangoBoost; Liangyu Zhao, University of Washington; Daniar Kurniawan, Chanmyeong Kim, Heetaek Jeong, Changsu Kim, and Hyeonseong Choi, MangoBoost; Liangcheng Yu, University of Pennsylvania; Arvind Krishnamurthy, University of Washington; Justine Sherry, Carnegie Mellon University; Eriko Nurvitadhi, MangoBoost
All-to-All(v) communication is a critical primitive in modern machine learning workloads, particularly mixture-of-experts (MoE) models. Unfortunately, efficient scheduling is challenging due to workload skew, heterogeneous two-tier fabrics, and incast congestion, compounded by the dynamic nature of MoE workloads, where traffic shifts every few hundred milliseconds. Existing schedulers are hardly scalable, incurring seconds to hours of synthesis time, making them impractical.
We present FAST, an efficient All-to-All(v) scheduler. FAST addresses skew through intra-server rebalancing and enforces balanced, one-to-one scale-out transfers that avoid incast. Evaluated extensively on both NVIDIA H200 and AMD MI300X clusters, FAST consistently outperforms state-of-the-art solutions on skewed workloads while reducing synthesis time by orders of magnitude.
View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
Yiran Lei, Carnegie Mellon University and MangoBoost; Dongjoo Lee, MangoBoost; Liangyu Zhao, University of Washington; Daniar Kurniawan, Chanmyeong Kim, Heetaek Jeong, Changsu Kim, and Hyeonseong Choi, MangoBoost; Liangcheng Yu, University of Pennsylvania; Arvind Krishnamurthy, University of Washington; Justine Sherry, Carnegie Mellon University; Eriko Nurvitadhi, MangoBoost
All-to-All(v) communication is a critical primitive in modern machine learning workloads, particularly mixture-of-experts (MoE) models. Unfortunately, efficient scheduling is challenging due to workload skew, heterogeneous two-tier fabrics, and incast congestion, compounded by the dynamic nature of MoE workloads, where traffic shifts every few hundred milliseconds. Existing schedulers are hardly scalable, incurring seconds to hours of synthesis time, making them impractical.
We present FAST, an efficient All-to-All(v) scheduler. FAST addresses skew through intra-server rebalancing and enforces balanced, one-to-one scale-out transfers that avoid incast. Evaluated extensively on both NVIDIA H200 and AMD MI300X clusters, FAST consistently outperforms state-of-the-art solutions on skewed workloads while reducing synthesis time by orders of magnitude.
View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions










