Uploaded August 2025 | Updated September 2026, 1 week ago
Running large language models at scale for 1 billion+ users is no small feat.
In "Inference Deployments and Comms Implication", Meta’s Cen Zhao, Xiaodong Wang, and Jianyu Huang will share how they’re tackling the compute and communication bottlenecks of LLM inference at production scale.
🔍 This session will dive deeper into topics like:
▪ Compute-bound vs. memory-bound stages in inference
▪ Context Parallelism & iRoPE for faster prefill
▪ Expert Parallelism + Tensor Parallelism optimizations
▪ Future directions: fused kernels, DDA, and more
If you’re working on scaling LLMs or optimizing inference systems, you won’t want to miss it! Register today: atscaleconference.com/events/scale-networking
Running large language models at scale for 1 billion+ users is no small feat.
In "Inference Deployments and Comms Implication", Meta’s Cen Zhao, Xiaodong Wang, and Jianyu Huang will share how they’re tackling the compute and communication bottlenecks of LLM inference at production scale.
🔍 This session will dive deeper into topics like:
▪ Compute-bound vs. memory-bound stages in inference
▪ Context Parallelism & iRoPE for faster prefill
▪ Expert Parallelism + Tensor Parallelism optimizations
▪ Future directions: fused kernels, DDA, and more
If you’re working on scaling LLMs or optimizing inference systems, you won’t want to miss it! Register today: atscaleconference.com/events/scale-networking










