Ray + vLLM  Efficient Multi Node Orchestration for Sparse MoE Model Serving | Ray Summit 2025 @anyscale
Ray + vLLM  Efficient Multi Node Orchestration for Sparse MoE Model Serving | Ray Summit 2025  @anyscale
Uploaded November 2025 | Updated September 2026, 2 weeks ago
Slides: drive.google.com/file/d/11OSdPJLZ1v4QH2KHlEYGYCts5qEdR5gN/view?usp=sharing

At Ray Summit 2025, Kourosh Hakhamaneshi and Seiji Eicher from Anyscale share how Ray Serve and vLLM are enabling efficient, scalable serving of Mixture-of-Experts (MoE) models—despite the complex orchestration challenges these architectures introduce.

They begin by outlining why MoE models are so appealing: selective expert activation offers cost-effective scaling for large language models. But to realize these efficiency gains in production, serving systems must handle very large batch sizes, where KV-cache memory quickly becomes the bottleneck.

The speakers dive into how Multi-head Latent Attention (MLA) helps reduce KV-cache footprint through low-rank compression—yet creates new challenges when combined with high degrees of expert parallelism (EP). When tensor parallelism is introduced, KV-cache duplication becomes unavoidable, often making data-parallel attention the better choice for MoE inference.

Kourosh and Seiji also highlight how combining data parallelism with expert parallelism unlocks unique optimizations for the prefill vs. decode phases of inference, making prefill/decode disaggregation a powerful strategy for maximizing utilization across heterogeneous resources.

These interconnected tradeoffs require sophisticated multi-node orchestration, and this is where Ray Serve + vLLM shine. The speakers show how the two systems work together to balance flexibility, high throughput, and operational simplicity—enabling practical, production-grade MoE serving at scale.

Attendees will learn how to design and operate distributed MoE inference pipelines, optimize KV-cache usage, and leverage Ray Serve to coordinate complex parallelism strategies across large GPU clusters.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
Ray + vLLM  Efficient Multi Node Orchestration for Sparse MoE Model Serving | Ray Summit 2025Hybrid RL + Imitation Learning for Robotics with Ray at RAI InstituteHow Runhouse Orchestrates Multi-Cluster Ray Workloads | Ray Summit 2025How vLLM and Ray Work TogetherCoinbases ML Training Evolution: From Sagemaker to Ray | Ray Summit 2024Secure & Scalable AI on Ray + Kubernetes: Google’s Decoupled Agent Pattern | Ray Summit 2025vLLM TPU: A new unified-backend supporting Pytorch and JAX natively on TPU | Ray Summit 2025Hugging Face + vLLM: One Model Definition to Rule Them All | Ray Summit 2025Matrix: Reliable Framework for Data-Centric Experimentation at Scale  | Ray Summit 2025From Spark to Ray: CSSs Data Revolution with Daft | Ray Summit 2024How Rubrik Unlocked AI at Scale with Ray Serve | Ray Summit 2024How Workday Achieved 50x Cheaper Model Serving with Ray Serve | Ray Summit 2025
Anyscale |

Ray + vLLM Efficient Multi Node Orchestration for Sparse MoE Model Serving | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER