NSDI 26 - FLARE: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus... @UsenixOrg
NSDI 26 - FLARE: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus...  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
FLARE: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale

Weihao Cui, Shanghai Jiao Tong University and National University of Singapore; Ji Zhang, Independent Researcher; Han Zhao, Shanghai Jiao Tong University; Chao Liu, Independent Researcher; Jian Sha, Tsinghua University; Bo Sang, Ant Group; Bingsheng He, National University of Singapore; Minyi Guo and Quan Chen, Shanghai Jiao Tong University

The rapid proliferation of large language models has driven the need for efficient GPU training clusters. However, it is challenging due to the frequent occurrence of training anomalies. Since existing diagnostic tools are narrowly tailored to specific issues, there are gaps in their ability to address anomalies spanning the entire training stack. In response, we introduce FLARE, a diagnostic framework designed for distributed LLM training at scale. FLARE first integrates a lightweight tracing daemon for full-stack and backend-extensible tracing. Additionally, it features a diagnostic engine that automatically diagnoses anomalies, with a focus on performance regressions. The deployment of FLARE across 6,000 GPUs has demonstrated significant improvements in pinpointing deficiencies in real-world scenarios, with continuous operation for over eight months.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - FLARE: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus...NSDI 26 - Remote TCP Connection Offload and ApplicationsUSENIX Security 25 - GraphAce: Secure Two-Party Graph Analysis Achieving Communication EfficiencyOpen Access at USENIXUSENIX Security 24 - Large Language Models for Code Analysis: Do LLMs Really Do Their Job?NSDI 26 - Cost-effective and Reliable Global Internet Peering with Programmable SwitchesSREcon26 Americas - Intelligent Load Balancing in KubernetesNSDI 26 - Mortise: Auto-tuning Congestion Control to Optimize QoE via Network-Aware ParameterNSDI 26 - Latency-Aware Caching with Delayed Hits: From Bursty Traffic to Pipeline ArchitecturesNSDI 26 - HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public CloudsNSDI 26 - Medley: Optimizing Midgress Bandwidth for Commercial Live Streaming CDNsNSDI 26 - Checkmate: Zero Performance Overhead Model Checkpointing via Network Gradient Replication
USENIX |

NSDI '26 - FLARE: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus...

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER