Dynamic Scheduling for Large Language Model Serving | Ray Summit 2024 @anyscale
Dynamic Scheduling for Large Language Model Serving | Ray Summit 2024  @anyscale
Uploaded October 2024 | Updated September 2026, 2 weeks ago
Hanyu Zhao from Alibaba Group presents Llumnix, a dynamic request scheduling system for large language models, at Ray Summit 2024. Built on vLLM and Ray, Llumnix addresses key challenges in LLM serving through innovative runtime rescheduling and KV cache migration across instances.

Zhao discusses how Llumnix reduces prefill latencies through cross-instance defragmentation and minimizes tail decoding latencies by balancing loads and reducing preemptions. The talk covers the research journey behind Llumnix, from its origins to its publication at OSDI '24, and its subsequent deployment and evolution at Alibaba.

The presentation provides insights into the current state of Llumnix and outlines future development plans. Zhao also highlights the open-source nature of the project, available on GitHub, encouraging community engagement and collaboration.

This session offers valuable information for those interested in optimizing LLM serving, particularly in large-scale, high-performance environments. It demonstrates practical applications of Ray and vLLM in addressing complex scheduling challenges in AI infrastructure.

--

Interested in more?
- Watch the full Day 1 Keynote: youtu.be/jwZHJthQvXo
- Watch the full Day 2 Keynote youtu.be/Lury2ad6KG8

--

đź”— Connect with us:
- Subscribe to our YouTube channel: youtube.com/@anyscale
- Twitter: https://x.com/anyscalecompute
- LinkedIn: linkedin.com/company/joinanyscale
- Website: anyscale.com
Dynamic Scheduling for Large Language Model Serving | Ray Summit 2024How Prime Intellect Builds Scalable Infrastructure for Agentic RL | Ray Summit 2025BentoML or RayServe, You Can Choose Both with BentoRayHow KubeRay Is Evolving for Massive AI Workloads  | Ray Summit 2025RDMA P2P Deep Dive: KvCache Transfer, Weight Updates & MoE Routing at Perplexity | Ray Summit 2025Ray Direct Transport: RDMA Support in Ray Core | Ray Summit 2025Enabling End-to-End LLMOps on Michelangelo with RayBuilding LLaMA: Metas Director of GenAI Sergey Edunov | Ray Summit 2024Maximizing Compute Efficiency on Anyscale | Ray Summit 2025The PARK Stack: The LAMP Stack of the AI Era | Ben LoricaAgentic Workload Inference at Scale: ByteDance’s AIBrix & DeerFlow | Ray Summit 2025Productionizing LLMs at Scale for Real-World Commerce | Ray on the Road – NYC 2025
Anyscale |

Dynamic Scheduling for Large Language Model Serving | Ray Summit 2024

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER