Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Siddharth Kodwani and Rong Zhou from Zoox share how they redesigned their ML model-serving infrastructure using Ray Serve—dramatically improving deployment velocity while strengthening system reliability across diverse production workloads.
They begin by outlining a key challenge faced by modern ML organizations: balancing rapid iteration with robust, fault-tolerant deployments. To solve this, the Zoox team built a self-service deployment platform powered by dynamic Ray Serve APIs. Their architecture introduces dedicated cluster isolation for different use cases, ensuring strong reliability guarantees while still enabling fast, independent deployments by engineering teams.
The speakers then walk through the platform’s built-in reliability checks, which automatically validate deployments in a self-serve workflow—reducing operations overhead while maintaining system integrity.
A major highlight is the seamless integration of LLMs and MLLMs via Ray LLM APIs. This capability has significantly boosted experimentation velocity, making it easy to onboard, benchmark, and iterate on new foundation models. The platform also supports cost-efficient batch inference, offering an economical alternative to third-party model APIs.
They share key results achieved through this architecture:
Automated self-serve deployments: Reduced deployment time from days to minutes with built-in reliability guardrails
Enhanced reliability: Dedicated clusters ensure performance isolation and predictable behavior
Accelerated LLM adoption: Streamlined onboarding enables rapid experimentation with LLM/MLLM models
Scalable, unified architecture: One platform that supports both traditional ML and large-scale LLM workloads in production
Attendees will learn how Zoox built a fast, resilient, and scalable model-serving platform—and how Ray Serve can power next-generation ML and LLM deployments across organizations.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
At Ray Summit 2025, Siddharth Kodwani and Rong Zhou from Zoox share how they redesigned their ML model-serving infrastructure using Ray Serve—dramatically improving deployment velocity while strengthening system reliability across diverse production workloads.
They begin by outlining a key challenge faced by modern ML organizations: balancing rapid iteration with robust, fault-tolerant deployments. To solve this, the Zoox team built a self-service deployment platform powered by dynamic Ray Serve APIs. Their architecture introduces dedicated cluster isolation for different use cases, ensuring strong reliability guarantees while still enabling fast, independent deployments by engineering teams.
The speakers then walk through the platform’s built-in reliability checks, which automatically validate deployments in a self-serve workflow—reducing operations overhead while maintaining system integrity.
A major highlight is the seamless integration of LLMs and MLLMs via Ray LLM APIs. This capability has significantly boosted experimentation velocity, making it easy to onboard, benchmark, and iterate on new foundation models. The platform also supports cost-efficient batch inference, offering an economical alternative to third-party model APIs.
They share key results achieved through this architecture:
Automated self-serve deployments: Reduced deployment time from days to minutes with built-in reliability guardrails
Enhanced reliability: Dedicated clusters ensure performance isolation and predictable behavior
Accelerated LLM adoption: Streamlined onboarding enables rapid experimentation with LLM/MLLM models
Scalable, unified architecture: One platform that supports both traditional ML and large-scale LLM workloads in production
Attendees will learn how Zoox built a fast, resilient, and scalable model-serving platform—and how Ray Serve can power next-generation ML and LLM deployments across organizations.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com










