Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Josh Karpel from Workday shares how the company rebuilt its ML model-serving architecture using Ray Serve—enabling massive scale, high availability, and dramatic cost reductions for Workday’s customers.
He begins by outlining the core challenge Workday faced in early 2023: serving dedicated ML models for every tenant across every environment had become increasingly expensive and difficult to scale. The solution was a ground-up redesign built on Ray Serve, which now powers tens of thousands of models across more than a dozen environments. With Ray Serve’s built-in autoscaling and efficient request routing, Workday achieved 50× cost savings while maintaining low latency and strong reliability.
Josh then covers the less conventional aspects of Workday’s system—unique usage patterns that pushed Ray Serve far beyond its original design. Early deployments ran into scalability ceilings with just a few dozen applications, but thanks to Ray’s open-source nature, Workday contributed a series of critical improvements that now allow Ray Serve to support thousands of applications per cluster.
In this talk, Josh dives deep into how Workday uses Ray Serve for large-scale model serving, the architectural challenges encountered while scaling, and the contributions made back to the Ray community to overcome them. Attendees will leave with a deeper understanding of Ray Serve internals, practical patterns for building complex serving systems, and inspiration to contribute to the ecosystem themselves.
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
At Ray Summit 2025, Josh Karpel from Workday shares how the company rebuilt its ML model-serving architecture using Ray Serve—enabling massive scale, high availability, and dramatic cost reductions for Workday’s customers.
He begins by outlining the core challenge Workday faced in early 2023: serving dedicated ML models for every tenant across every environment had become increasingly expensive and difficult to scale. The solution was a ground-up redesign built on Ray Serve, which now powers tens of thousands of models across more than a dozen environments. With Ray Serve’s built-in autoscaling and efficient request routing, Workday achieved 50× cost savings while maintaining low latency and strong reliability.
Josh then covers the less conventional aspects of Workday’s system—unique usage patterns that pushed Ray Serve far beyond its original design. Early deployments ran into scalability ceilings with just a few dozen applications, but thanks to Ray’s open-source nature, Workday contributed a series of critical improvements that now allow Ray Serve to support thousands of applications per cluster.
In this talk, Josh dives deep into how Workday uses Ray Serve for large-scale model serving, the architectural challenges encountered while scaling, and the contributions made back to the Ray community to overcome them. Attendees will leave with a deeper understanding of Ray Serve internals, practical patterns for building complex serving systems, and inspiration to contribute to the ecosystem themselves.
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com










