Enhancing Runtime Reliability in LLM Training via Fine-Grained Observability - Live from SCC @scaleconference
Enhancing Runtime Reliability in LLM Training via Fine-Grained Observability - Live from SCC  @scaleconference
Uploaded August 2025 | Updated September 2026, 1 week ago
Speaker: Lei Zhang from ByteDance

Learn more here: atscaleconference.com/events/scale-networking
Enhancing Runtime Reliability in LLM Training via Fine-Grained Observability - Live from SCCPhysical Network Design at Scale | Brandon Premo and Richard Cziva from MetaPerformance Optimizations at 100K+ Scale - Live from SCCSecuring Production Debugging at Hyperscale | Shridivya Sharma and Luke Kelly from MicrosoftPyTorch Symmetric Memory: A New Paradigm for Programming Distributed AI - Live from SCCScaling is one thing. Keeping it running is another.Scaling AI Network with DSF by Ron He and Ankur SinghHow Meta Avoided a Major OutageBridging the Intent Gap in Agentic Systems | Anoop Deoras from AWS (Live)Lightning Talk: AI-Native Network Operation | YuLing Chen, Meta@Scale Podcast with Founder of Juniper Networks, Dr. Pradeep SindhuLive from SCCC: Virgo: Scale-out Data Center Network for AI | Jeongkeun JK Lee from Google
@Scale |

Enhancing Runtime Reliability in LLM Training via Fine-Grained Observability - Live from SCC

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER