Uploaded November 2025 | Updated September 2026, 1 week ago
At Ray Summit 2025, Yongji Wu from UC Berkeley and Rui Qiao from Anyscale share how they are advancing large-scale Expert Parallelism (EP) to unlock efficient, scalable inference for Mixture-of-Experts (MoE) models.
They begin by outlining a core constraint in MoE serving: EP often requires massive, monolithic deployment units—for example, DeepSeek R3/V1 needs 144 GPUs just to form a single serving instance. Such large units make it extremely difficult for traditional inter-instance autoscaling systems to react to real-world workload fluctuations.
To address this, the speakers introduce intra-instance Elastic EP, a new technique that brings fine-grained, low-latency autoscaling inside a single EP instance. This enables vLLM to tightly match GPU resources to workload demand without incurring downtime, fragmentation, or inefficient overprovisioning.
They then show how Ray is used to orchestrate Elastic EP scaling across distributed clusters, providing the coordination, lifecycle management, and flexibility needed to dynamically adjust expert-parallel resources while keeping inference fast and reliable.
Attendees will learn practical strategies for serving large MoE models at scale, optimizing KV-cache and expert utilization, and using Ray to coordinate sophisticated intra-instance parallelism patterns.
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
đź”— Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
At Ray Summit 2025, Yongji Wu from UC Berkeley and Rui Qiao from Anyscale share how they are advancing large-scale Expert Parallelism (EP) to unlock efficient, scalable inference for Mixture-of-Experts (MoE) models.
They begin by outlining a core constraint in MoE serving: EP often requires massive, monolithic deployment units—for example, DeepSeek R3/V1 needs 144 GPUs just to form a single serving instance. Such large units make it extremely difficult for traditional inter-instance autoscaling systems to react to real-world workload fluctuations.
To address this, the speakers introduce intra-instance Elastic EP, a new technique that brings fine-grained, low-latency autoscaling inside a single EP instance. This enables vLLM to tightly match GPU resources to workload demand without incurring downtime, fragmentation, or inefficient overprovisioning.
They then show how Ray is used to orchestrate Elastic EP scaling across distributed clusters, providing the coordination, lifecycle management, and flexibility needed to dynamically adjust expert-parallel resources while keeping inference fast and reliable.
Attendees will learn practical strategies for serving large MoE models at scale, optimizing KV-cache and expert utilization, and using Ray to coordinate sophisticated intra-instance parallelism patterns.
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
đź”— Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com










