SGLang: An Efficient Open-Source Framework for Large-Scale LLM Serving | Ray Summit 2025 @anyscale
SGLang: An Efficient Open-Source Framework for Large-Scale LLM Serving | Ray Summit 2025  @anyscale
Uploaded November 2025 | Updated September 2026, 1 week ago
At Ray Summit 2025, Ying Sheng from SGLang and Qiaolin Yu from Anyscale share how SGLang has become a high-performance framework for LLM serving—powering production workloads at leading companies through its optimized architecture and next-generation inference capabilities.

They begin by breaking down SGLang’s core features, including its lightweight execution engine, optimized KVCache handling, and flexible serving abstractions designed for both high-throughput batch workloads and ultra–low-latency interactive use cases.

The speakers then dive into the key performance optimization techniques that enable SGLang to consistently outperform traditional serving stacks. Topics include scheduling strategies, memory management improvements, parallel execution paths, and advanced kernel optimizations tailored for modern accelerators.

Ying and Qiaolin also share insights from real-world production deployments, highlighting how companies use SGLang to support rapid model iteration, cost-efficient scaling, and stable high-volume traffic patterns.

Finally, they present the future roadmap for SGLang—covering upcoming performance enhancements, deeper hardware integration, and new features aimed at simplifying large-scale LLM serving for both enterprises and open-source users.

Attendees will leave with a deep understanding of how SGLang achieves best-in-class performance and where the framework is headed next.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
SGLang: An Efficient Open-Source Framework for Large-Scale LLM Serving | Ray Summit 2025CoServe: Max Performance, Minimal Compute | Ray Summit 2025How Ray Data Powers Scalable AI Workloads | Ray Summit 2025Inside Uber: Scaling Model Training with Ray | Ray Summit 2025Transforming Multimodal Data Management with LanceDB-Ray | Ray Summit 2024Why Ray Became a Distributed Computing Engine for Modern AIRLlib: Lessons from the V2 Stack and Road Ahead | Ray Summit 2025Pinterests ML Evolution: Distributed Training with Ray | Ray Summit 2024How xAI Scales Image & Video Processing with Ray | Ray Summit 2025Ion Stoica on Agentic Systems and AI Reliability | Ray on the Road – NYC 2025Anyscale on Azure: Build and deploy AI at scale in your own tenantAn Overview of CloudKitchenss Ray-Powered ML Platform | Ray Summit 2024
Anyscale |

SGLang: An Efficient Open-Source Framework for Large-Scale LLM Serving | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER