Uploaded June 2026 | Updated September 2026, 1 week ago
Abstract
For a decade, data center fabric was the boring part of the stack. CLOS topologies, ECMP, and 100/400G optics mostly just worked for the cloud and CDN workloads that dominated the 2010s. Then AI training arrived. A synchronous collective across tens of thousands of GPUs is gated by its slowest participant, and the fabric (long presumed healthy if every link was up and every BGP session was stable) became the throughput ceiling.
This talk is about the class of failures operators are now burning hours on: gray failures. Links that are nominally up but quietly corrupting. Optics degrading through rising FEC correction counts before they ever go down. ECMP hash polarization that lands an all-reduce on a single hot path. PFC headroom exhaustion that never trips a threshold. Microbursts invisible to any 30-second SNMP poll. None of these show up on the NOC wall. All of them show up on the compute bill.
Drawing on production examples from AI fabrics, this session walks through what gray failures look like in real telemetry, why standard tooling (polled SNMP, sampled flow, ping/iperf) systematically misses them, and which techniques (high-resolution synchronized timestamping, tail-latency percentiles, streaming telemetry, host-side straggler correlation) catch them before a training job silently loses throughput across a multi-week run.
The operations skills this community already has are load-bearing again. This talk is about what to re-tool, and what to keep.
nanog.org/events/nanog-97/content/5774
Abstract
For a decade, data center fabric was the boring part of the stack. CLOS topologies, ECMP, and 100/400G optics mostly just worked for the cloud and CDN workloads that dominated the 2010s. Then AI training arrived. A synchronous collective across tens of thousands of GPUs is gated by its slowest participant, and the fabric (long presumed healthy if every link was up and every BGP session was stable) became the throughput ceiling.
This talk is about the class of failures operators are now burning hours on: gray failures. Links that are nominally up but quietly corrupting. Optics degrading through rising FEC correction counts before they ever go down. ECMP hash polarization that lands an all-reduce on a single hot path. PFC headroom exhaustion that never trips a threshold. Microbursts invisible to any 30-second SNMP poll. None of these show up on the NOC wall. All of them show up on the compute bill.
Drawing on production examples from AI fabrics, this session walks through what gray failures look like in real telemetry, why standard tooling (polled SNMP, sampled flow, ping/iperf) systematically misses them, and which techniques (high-resolution synchronized timestamping, tail-latency percentiles, streaming telemetry, host-side straggler correlation) catch them before a training job silently loses throughput across a multi-week run.
The operations skills this community already has are load-bearing again. This talk is about what to re-tool, and what to keep.
nanog.org/events/nanog-97/content/5774










