Uploaded June 2025 | Updated September 2026, 2 weeks ago
Running Ray in production unlocks powerful performance and scalability—but it also comes with real-world operational challenges.
Watch this webinar to learn more about:
- Reference architectures for running Ray in production
- Best practices for improving stability and observability
- Common infrastructure challenges and how to address them
- How RayTurbo accelerates performance and scaling in production environments
Running Ray in production unlocks powerful performance and scalability—but it also comes with real-world operational challenges.
Watch this webinar to learn more about:
- Reference architectures for running Ray in production
- Best practices for improving stability and observability
- Common infrastructure challenges and how to address them
- How RayTurbo accelerates performance and scaling in production environments








![[Ray Meetup] Ray + vLLM in Action: Lessons from Pinterest and Large Scale Distributed Inference
Listen in to our Ray Meetup where we explored batch inference at scale with Ray and vLLM! Learn how Pinterest scales batch inference using Ray, and get a first look at Anyscale’s latest tools—Ray Serve and Data LLM—for orchestrating large-scale LLM inference. We’ll cover topics like batch inference, prefill-decode disaggregation, DP/EP parallelism, and custom request routing.
Speakers:
Chia-Wei Chen, Software Engineer, ML Training Infra, Pinterest
Kourosh Hakhamaneshi, AI Lead, Anyscale [Ray Meetup] Ray + vLLM in Action: Lessons from Pinterest and Large Scale Distributed Inference](https://i.ytimg.com/vi/HDSy09hrm2I/mqdefault.jpg)

