[Ray Meetup] Ray + vLLM in Action: Lessons from Pinterest and Large Scale Distributed Inference @anyscale
[Ray Meetup] Ray + vLLM in Action: Lessons from Pinterest and Large Scale Distributed Inference  @anyscale
Uploaded June 2025 | Updated September 2026, 2 weeks ago
Listen in to our Ray Meetup where we explored batch inference at scale with Ray and vLLM! Learn how Pinterest scales batch inference using Ray, and get a first look at Anyscale’s latest tools—Ray Serve and Data LLM—for orchestrating large-scale LLM inference. We’ll cover topics like batch inference, prefill-decode disaggregation, DP/EP parallelism, and custom request routing.

Speakers:
Chia-Wei Chen, Software Engineer, ML Training Infra, Pinterest
Kourosh Hakhamaneshi, AI Lead, Anyscale
[Ray Meetup] Ray + vLLM in Action: Lessons from Pinterest and Large Scale Distributed InferenceHow Zoox Built a Reliable, High-Velocity Model Serving Platform with Ray Serve | Ray Summit 2025Scaling Machine Learning at Tripadvisor: Our Journey with Ray and Anyscale | Ray Summit 2025AWS + vLLM: Building the Future of Open, Fast LLM Serving | Ray Summit 2025Ray + vLLM  Efficient Multi Node Orchestration for Sparse MoE Model Serving | Ray Summit 2025Hybrid RL + Imitation Learning for Robotics with Ray at RAI InstituteHow Runhouse Orchestrates Multi-Cluster Ray Workloads | Ray Summit 2025How vLLM and Ray Work TogetherCoinbases ML Training Evolution: From Sagemaker to Ray | Ray Summit 2024Secure & Scalable AI on Ray + Kubernetes: Google’s Decoupled Agent Pattern | Ray Summit 2025vLLM TPU: A new unified-backend supporting Pytorch and JAX natively on TPU | Ray Summit 2025Hugging Face + vLLM: One Model Definition to Rule Them All | Ray Summit 2025
Anyscale |

[Ray Meetup] Ray + vLLM in Action: Lessons from Pinterest and Large Scale Distributed Inference

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER