Elastic Expert Parallelism for vLLM | Ray Summit 2025 @anyscale
Elastic Expert Parallelism for vLLM | Ray Summit 2025  @anyscale
Uploaded November 2025 | Updated September 2026, 1 week ago
At Ray Summit 2025, Yongji Wu from UC Berkeley and Rui Qiao from Anyscale share how they are advancing large-scale Expert Parallelism (EP) to unlock efficient, scalable inference for Mixture-of-Experts (MoE) models.

They begin by outlining a core constraint in MoE serving: EP often requires massive, monolithic deployment units—for example, DeepSeek R3/V1 needs 144 GPUs just to form a single serving instance. Such large units make it extremely difficult for traditional inter-instance autoscaling systems to react to real-world workload fluctuations.

To address this, the speakers introduce intra-instance Elastic EP, a new technique that brings fine-grained, low-latency autoscaling inside a single EP instance. This enables vLLM to tightly match GPU resources to workload demand without incurring downtime, fragmentation, or inefficient overprovisioning.

They then show how Ray is used to orchestrate Elastic EP scaling across distributed clusters, providing the coordination, lifecycle management, and flexibility needed to dynamically adjust expert-parallel resources while keeping inference fast and reliable.

Attendees will learn practical strategies for serving large MoE models at scale, optimizing KV-cache and expert utilization, and using Ray to coordinate sophisticated intra-instance parallelism patterns.

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

đź”— Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
Elastic Expert Parallelism for vLLM | Ray Summit 2025Apple’s Approach to Scalable Machine Learning Infrastructure on Ray | Ray Summit 2025How Uber Optimize Marketplaces with Ray | Ray Summit 2024Improved Scheduling Flexibility with Label Selectors in Ray | Ray Summit 2025Fighting Fire with Algorithms: Lockheeds RL-Based Wildfire Solution | Ray Summit 2024Meet verl: An RL Framework for LLM Reasoning & Tool Use | Ray Summit 2025How BMW Scales Automotive AI Workloads with the Ray Framework | Ray Summit 2025Introduction to Anyscale and Ray AI LibrariesState of vLLM 2025 | Ray Summit 2025Intelligent Data Classification with Ray and vLLM at Apple | Ray Summit 2024Ray Libraries in Practice: Multimodal AI WorkloadsLiquidAI’s Approach to Large-Scale Synthetic Data Generation Using Ray | Ray Summit 2025
Anyscale |

Elastic Expert Parallelism for vLLM | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER