Uploaded August 2025 | Updated September 2026, 1 week ago
As large language model (LLM) training scales across tens of thousands of GPUs, ensuring runtime reliability becomes both more challenging and more critical for maintaining efficiency. This talk explores how fine-grained observability can substantially enhance reliability in LLM training at scale. First, we discuss automated methods for detecting faulty machines by leveraging distinctive monitoring metric patterns, enabling rapid and accurate identification of problematic nodes while minimizing manual intervention.
Second, we tackle reliability challenges within collective communication libraries (CCL), introducing a lightweight tracing and root cause analysis system that treats CCL as system software and reveals internal control and data dependencies. This approach allows for swift and precise detection of communication-related anomalies.
Collectively, these advancements illustrate how fine-grained observability at both the machine and communication levels can significantly improve the robustness and operational efficiency of large-scale LLM training.
Speaker: Lei Zhang from ByteDance
Learn more here: atscaleconference.com
As large language model (LLM) training scales across tens of thousands of GPUs, ensuring runtime reliability becomes both more challenging and more critical for maintaining efficiency. This talk explores how fine-grained observability can substantially enhance reliability in LLM training at scale. First, we discuss automated methods for detecting faulty machines by leveraging distinctive monitoring metric patterns, enabling rapid and accurate identification of problematic nodes while minimizing manual intervention.
Second, we tackle reliability challenges within collective communication libraries (CCL), introducing a lightweight tracing and root cause analysis system that treats CCL as system software and reveals internal control and data dependencies. This approach allows for swift and precise detection of communication-related anomalies.
Collectively, these advancements illustrate how fine-grained observability at both the machine and communication levels can significantly improve the robustness and operational efficiency of large-scale LLM training.
Speaker: Lei Zhang from ByteDance
Learn more here: atscaleconference.com








![Evolving GenAI Media Infrastructure Deployments | Rushaan Mahajan, Sima Labs
This presentation by Rushaan Mahajan from Sima Labs explores how to optimize video generation infrastructure to meet the growing demand for personalized, high-quality video experiences.
Video generation is a computationally intensive process that requires significantly more resources than text generation, leading to high latency and costs. [00:35]
The key challenges in video inference include the iterative denoising loop, the memory and bandwidth bottleneck in the VAE decoder, and the need to optimize the entire runtime stack to achieve real-time, high-fidelity, and personalized video generation at scale.
Optimizing the balance between the VAE compression ratio and the denoiser complexity is crucial to reducing the overall computational cost of the video generation pipeline. [09:52]
Techniques like latent space compression, sampling optimization, caching, and pruning can significantly improve the runtime efficiency of video diffusion models without compromising quality. [11:39]
A multi-GPU strategy that utilizes different GPU types for different tasks (base generation, super-resolution, personalization) can help scale video generation without proportional cost increases. [13:35]
The goal is to make personalized video generation truly instant, transitioning it from a compute-bound novelty to a mainstream, interactive, and globally relevant platform. [13:57] Evolving GenAI Media Infrastructure Deployments | Rushaan Mahajan, Sima Labs](https://i.ytimg.com/vi/PrZpl4w1Lxk/mqdefault.jpg)

