Taming AI Infrastructure Failures with Agentic Debugging | Phillip Liu from Meta @scaleconference
Taming AI Infrastructure Failures with Agentic Debugging | Phillip Liu from Meta  @scaleconference
Uploaded June 2026 | Updated September 2026, 1 week ago
NCCL watchdog timeouts are a common failure mode in distributed AI model training. They impact not only Meta, but broadly affect anyone running PyTorch distributed training—and they’re notoriously hard to debug: even experts can spend hours triaging a single incident, and non-experts may be unable to root-cause them at all.

Over the past year, we investigated NCCL watchdog timeouts internally at Meta and partnered with the PyTorch community to categorize the major root-cause buckets. We then distilled these learnings into a practical decision tree and runbook to speed up triage and make debugging more accessible. We also explored using agent-based approaches to assist root-cause analysis and saw strong early results.

Learn more about the @Scale conferences here: atscaleconference.com
Taming AI Infrastructure Failures with Agentic Debugging | Phillip Liu from Meta@Scale Podcast with Founder of Juniper Networks, Dr. Pradeep SindhuTransparent MultiNIC Routing for Large AI Models - Live from SCCMeta’s DC Networks for Generative AI by Rohit Puri and Hany MorsyClosing Keynote with Dr. Pradeep Sindhu from Microsoft and Juniper NetworksWhy Your AI Voice Assistant Keeps Interrupting You (And How We Fixed It)Agentic Data at Scale: Transforming Data Experiences at Meta | Dinkar Pataballa from MetaGoing Behind the Scenes of Meta AI Glasses PrivacyEver wonder who REALLY created that viral video you just shared?Call for Speakers: @Scale Systems & ReliabilitySystems & Reliability 2026 - Recap!Invisible Watermarking: Content Provenance for Videos at Scale | Wes Castro, Meta
@Scale |

Taming AI Infrastructure Failures with Agentic Debugging | Phillip Liu from Meta

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER