AWS + vLLM: Building the Future of Open, Fast LLM Serving | Ray Summit 2025 @anyscale
AWS + vLLM: Building the Future of Open, Fast LLM Serving | Ray Summit 2025  @anyscale
Uploaded December 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Phi Nguyen from AWS shares how Amazon is advancing large-scale LLM inference through deep support and contributions to vLLM, the leading open-source engine for high-throughput, low-latency model serving.
He begins by highlighting how vLLM is used broadly across Amazon—including as a foundational component of the Amazon Rufus shopping assistant, which serves millions of customers. vLLM’s robust support for heterogeneous hardware, such as AWS Trainium and NVIDIA GPUs, enables Amazon to deploy a cost-optimized, multi-node inference architecture that routes requests to the most appropriate accelerator. This hybrid design delivers substantial cost savings while maintaining top-tier performance.

Phi then walks through:

Deployment best practices for running vLLM on AWS at scale
How Amazon builds multi-accelerator inference clusters using Trainium and GPUs
Open-source work streams and contributions Amazon has made to vLLM
Additional initiatives aimed at strengthening the vLLM ecosystem for AWS customers and the broader community

Attendees will gain insight into how Amazon runs vLLM in production at massive scale, how to architect heterogeneous inference pipelines on AWS, and how Amazon is helping drive the future of open-source LLM serving.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
AWS + vLLM: Building the Future of Open, Fast LLM Serving | Ray Summit 2025Ray + vLLM  Efficient Multi Node Orchestration for Sparse MoE Model Serving | Ray Summit 2025Hybrid RL + Imitation Learning for Robotics with Ray at RAI InstituteHow Runhouse Orchestrates Multi-Cluster Ray Workloads | Ray Summit 2025How vLLM and Ray Work TogetherCoinbases ML Training Evolution: From Sagemaker to Ray | Ray Summit 2024Secure & Scalable AI on Ray + Kubernetes: Google’s Decoupled Agent Pattern | Ray Summit 2025vLLM TPU: A new unified-backend supporting Pytorch and JAX natively on TPU | Ray Summit 2025Hugging Face + vLLM: One Model Definition to Rule Them All | Ray Summit 2025Matrix: Reliable Framework for Data-Centric Experimentation at Scale  | Ray Summit 2025From Spark to Ray: CSSs Data Revolution with Daft | Ray Summit 2024How Rubrik Unlocked AI at Scale with Ray Serve | Ray Summit 2024
Anyscale |

AWS + vLLM: Building the Future of Open, Fast LLM Serving | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER