Uploaded December 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Phi Nguyen from AWS shares how Amazon is advancing large-scale LLM inference through deep support and contributions to vLLM, the leading open-source engine for high-throughput, low-latency model serving.
He begins by highlighting how vLLM is used broadly across Amazon—including as a foundational component of the Amazon Rufus shopping assistant, which serves millions of customers. vLLM’s robust support for heterogeneous hardware, such as AWS Trainium and NVIDIA GPUs, enables Amazon to deploy a cost-optimized, multi-node inference architecture that routes requests to the most appropriate accelerator. This hybrid design delivers substantial cost savings while maintaining top-tier performance.
Phi then walks through:
Deployment best practices for running vLLM on AWS at scale
How Amazon builds multi-accelerator inference clusters using Trainium and GPUs
Open-source work streams and contributions Amazon has made to vLLM
Additional initiatives aimed at strengthening the vLLM ecosystem for AWS customers and the broader community
Attendees will gain insight into how Amazon runs vLLM in production at massive scale, how to architect heterogeneous inference pipelines on AWS, and how Amazon is helping drive the future of open-source LLM serving.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
At Ray Summit 2025, Phi Nguyen from AWS shares how Amazon is advancing large-scale LLM inference through deep support and contributions to vLLM, the leading open-source engine for high-throughput, low-latency model serving.
He begins by highlighting how vLLM is used broadly across Amazon—including as a foundational component of the Amazon Rufus shopping assistant, which serves millions of customers. vLLM’s robust support for heterogeneous hardware, such as AWS Trainium and NVIDIA GPUs, enables Amazon to deploy a cost-optimized, multi-node inference architecture that routes requests to the most appropriate accelerator. This hybrid design delivers substantial cost savings while maintaining top-tier performance.
Phi then walks through:
Deployment best practices for running vLLM on AWS at scale
How Amazon builds multi-accelerator inference clusters using Trainium and GPUs
Open-source work streams and contributions Amazon has made to vLLM
Additional initiatives aimed at strengthening the vLLM ecosystem for AWS customers and the broader community
Attendees will gain insight into how Amazon runs vLLM in production at massive scale, how to architect heterogeneous inference pipelines on AWS, and how Amazon is helping drive the future of open-source LLM serving.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale










