Inference Deployments and Comms Implication at @Scale: Networking 2025 @scaleconference
Inference Deployments and Comms Implication at @Scale: Networking 2025  @scaleconference
Uploaded August 2025 | Updated September 2026, 1 week ago
Running large language models at scale for 1 billion+ users is no small feat.

In "Inference Deployments and Comms Implication", Meta’s Cen Zhao, Xiaodong Wang, and Jianyu Huang will share how they’re tackling the compute and communication bottlenecks of LLM inference at production scale.

🔍 This session will dive deeper into topics like:
▪ Compute-bound vs. memory-bound stages in inference
▪ Context Parallelism & iRoPE for faster prefill
▪ Expert Parallelism + Tensor Parallelism optimizations
▪ Future directions: fused kernels, DDA, and more

If you’re working on scaling LLMs or optimizing inference systems, you won’t want to miss it! Register today: atscaleconference.com/events/scale-networking
Inference Deployments and Comms Implication at @Scale: Networking 2025Why Should You Attend?New scale. New challenges.Live Technology Panel from the SCC by Omar BaldonadoWhy do engineers keep coming back to @Scale?@Scale: Networking on August 13!How Meta Deployed Video Super Resolution at Scale | Ryan Lei, MetaLive from SCCC: Co-Designing Communication for AI Accelerators | Rajeev Nair and Wes BlandMaking Our Debut, in Bellevue!10x Backbone: Scaling Backbone Connectivity to Serve AI Demands at @Scale: Networking 2025AI is About to Change Video Calls Forever!Taming AI Infrastructure Failures with Agentic Debugging | Phillip Liu from Meta
@Scale |

"Inference Deployments and Comms Implication" at @Scale: Networking 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER