Enhancing Runtime Reliability in LLM Training via Fine-Grained Observability by Lei Zhang @scaleconference
Enhancing Runtime Reliability in LLM Training via Fine-Grained Observability by Lei Zhang  @scaleconference
Uploaded August 2025 | Updated September 2026, 1 week ago
As large language model (LLM) training scales across tens of thousands of GPUs, ensuring runtime reliability becomes both more challenging and more critical for maintaining efficiency. This talk explores how fine-grained observability can substantially enhance reliability in LLM training at scale. First, we discuss automated methods for detecting faulty machines by leveraging distinctive monitoring metric patterns, enabling rapid and accurate identification of problematic nodes while minimizing manual intervention.

Second, we tackle reliability challenges within collective communication libraries (CCL), introducing a lightweight tracing and root cause analysis system that treats CCL as system software and reveals internal control and data dependencies. This approach allows for swift and precise detection of communication-related anomalies.

Collectively, these advancements illustrate how fine-grained observability at both the machine and communication levels can significantly improve the robustness and operational efficiency of large-scale LLM training.

Speaker: Lei Zhang from ByteDance

Learn more here: atscaleconference.com
Enhancing Runtime Reliability in LLM Training via Fine-Grained Observability by Lei ZhangBuilding Storage for AI at ScaleLearn from Zoom’s AI Product Manager, Danran Chen!How Meta is Advancing Flash Storage with High-Density QLC Deployments!Live from SCCC: Fiber Networks Innovations for AI at Scale | Fabrice Ouandji and Sebastian GaultScaling Llama4 Training to 100K - Live from SCCThe Agentic Infrastructure Gap: In-Distribution Languages Make It a Coding Problem | Joe DuffyLightning Talk: AI-Native Network Capacity Management | Mohab Gawish, MetaLive from SCCC: MetaRoCE: From Spec to NIC to Open Source | Balakrishnan Raman and Sandeep NagarajEvolving GenAI Media Infrastructure Deployments | Rushaan Mahajan, Sima LabsOur Journey to Safely Unleash Agents at Meta Scale | David Pariag from MetaLive from SCCC: MetaRoCE: Meta’s RDMA Transport | Arvind Srinivasan and Kingshuk Mandal
@Scale |

Enhancing Runtime Reliability in LLM Training via Fine-Grained Observability by Lei Zhang

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER