How Zoox Built a Reliable, High-Velocity Model Serving Platform with Ray Serve | Ray Summit 2025 @anyscale
How Zoox Built a Reliable, High-Velocity Model Serving Platform with Ray Serve | Ray Summit 2025  @anyscale
Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Siddharth Kodwani and Rong Zhou from Zoox share how they redesigned their ML model-serving infrastructure using Ray Serve—dramatically improving deployment velocity while strengthening system reliability across diverse production workloads.

They begin by outlining a key challenge faced by modern ML organizations: balancing rapid iteration with robust, fault-tolerant deployments. To solve this, the Zoox team built a self-service deployment platform powered by dynamic Ray Serve APIs. Their architecture introduces dedicated cluster isolation for different use cases, ensuring strong reliability guarantees while still enabling fast, independent deployments by engineering teams.

The speakers then walk through the platform’s built-in reliability checks, which automatically validate deployments in a self-serve workflow—reducing operations overhead while maintaining system integrity.

A major highlight is the seamless integration of LLMs and MLLMs via Ray LLM APIs. This capability has significantly boosted experimentation velocity, making it easy to onboard, benchmark, and iterate on new foundation models. The platform also supports cost-efficient batch inference, offering an economical alternative to third-party model APIs.

They share key results achieved through this architecture:

Automated self-serve deployments: Reduced deployment time from days to minutes with built-in reliability guardrails

Enhanced reliability: Dedicated clusters ensure performance isolation and predictable behavior

Accelerated LLM adoption: Streamlined onboarding enables rapid experimentation with LLM/MLLM models

Scalable, unified architecture: One platform that supports both traditional ML and large-scale LLM workloads in production

Attendees will learn how Zoox built a fast, resilient, and scalable model-serving platform—and how Ray Serve can power next-generation ML and LLM deployments across organizations.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
How Zoox Built a Reliable, High-Velocity Model Serving Platform with Ray Serve | Ray Summit 2025Scaling Machine Learning at Tripadvisor: Our Journey with Ray and Anyscale | Ray Summit 2025AWS + vLLM: Building the Future of Open, Fast LLM Serving | Ray Summit 2025Ray + vLLM  Efficient Multi Node Orchestration for Sparse MoE Model Serving | Ray Summit 2025Hybrid RL + Imitation Learning for Robotics with Ray at RAI InstituteHow Runhouse Orchestrates Multi-Cluster Ray Workloads | Ray Summit 2025How vLLM and Ray Work TogetherCoinbases ML Training Evolution: From Sagemaker to Ray | Ray Summit 2024Secure & Scalable AI on Ray + Kubernetes: Google’s Decoupled Agent Pattern | Ray Summit 2025vLLM TPU: A new unified-backend supporting Pytorch and JAX natively on TPU | Ray Summit 2025Hugging Face + vLLM: One Model Definition to Rule Them All | Ray Summit 2025Matrix: Reliable Framework for Data-Centric Experimentation at Scale  | Ray Summit 2025
Anyscale |

How Zoox Built a Reliable, High-Velocity Model Serving Platform with Ray Serve | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER