Inference Deployments and Comms Implication by Cen Zhao, Xiaodong Wang, and Jianyu Huang @scaleconference
Inference Deployments and Comms Implication by Cen Zhao, Xiaodong Wang, and Jianyu Huang  @scaleconference
Uploaded August 2025 | Updated September 2026, 2 weeks ago
This talk addresses the challenges and solutions for scaling large language model (LLM) inference to support up to 1 billion monthly active users across platforms for Meta AI, focusing on compute-bound prefill and memory-bound decode stages. Key challenges include the quadratic scaling of attention operations with sequence length and the linear growth of the KV cache, along with network-intensive operations impacting latency.

To enhance scaling efficiency, a multi-dimensional parallelism strategy is proposed across various hardware platforms, including Nvidia and AMD. Innovations such as Context Parallelism (CP) and iRoPE enable near-linear prefill scaling, while optimized communication techniques like Dynamic/Persistent All-to-All for Expert Parallelism (EP) and Direct Data Access (DDA) for Tensor Parallelism (TP) significantly improve performance. Future efforts aim to further enhance system efficiency through fused kernels and device-initiated operations.

Learn more here: atscaleconference.com
Inference Deployments and Comms Implication by Cen Zhao, Xiaodong Wang, and Jianyu HuangRDMA at Cloud Scale: The OCI Experience at @Scale: Networking 2025One thing attendees love about @ Scale?Live Panel: Agentic Autonomy and Evolution of Software & ResearchKeynote from Microsoft - Live from SCCWhats all the hype about? And who is behind building it?What if AI could help run the network powering billions of users?Live from SCCC: Anatomy of @scale Machine Learning Network Operations | Jim JulsonBuilding Responsive AI Agents with Real-Time Communication | Blaise Thomas, AgoraAI Conversations That Actually Feel HumanMeta Keynote - Ime Archibong, VP of Product at MetaHow Meta Is Letting Users Tune Their Own Recommendations
@Scale |

Inference Deployments and Comms Implication by Cen Zhao, Xiaodong Wang, and Jianyu Huang

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER