How Workday Achieved 50x Cheaper Model Serving with Ray Serve | Ray Summit 2025 @anyscale
How Workday Achieved 50x Cheaper Model Serving with Ray Serve | Ray Summit 2025  @anyscale
Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Josh Karpel from Workday shares how the company rebuilt its ML model-serving architecture using Ray Serve—enabling massive scale, high availability, and dramatic cost reductions for Workday’s customers.

He begins by outlining the core challenge Workday faced in early 2023: serving dedicated ML models for every tenant across every environment had become increasingly expensive and difficult to scale. The solution was a ground-up redesign built on Ray Serve, which now powers tens of thousands of models across more than a dozen environments. With Ray Serve’s built-in autoscaling and efficient request routing, Workday achieved 50× cost savings while maintaining low latency and strong reliability.

Josh then covers the less conventional aspects of Workday’s system—unique usage patterns that pushed Ray Serve far beyond its original design. Early deployments ran into scalability ceilings with just a few dozen applications, but thanks to Ray’s open-source nature, Workday contributed a series of critical improvements that now allow Ray Serve to support thousands of applications per cluster.

In this talk, Josh dives deep into how Workday uses Ray Serve for large-scale model serving, the architectural challenges encountered while scaling, and the contributions made back to the Ray community to overcome them. Attendees will leave with a deeper understanding of Ray Serve internals, practical patterns for building complex serving systems, and inspiration to contribute to the ecosystem themselves.

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
How Workday Achieved 50x Cheaper Model Serving with Ray Serve | Ray Summit 2025ByteDances Platform for Reinforcement Learning from Human Feedback | Ray Summit 2024Dynamic Scheduling for Large Language Model Serving | Ray Summit 2024How Prime Intellect Builds Scalable Infrastructure for Agentic RL | Ray Summit 2025BentoML or RayServe, You Can Choose Both with BentoRayHow KubeRay Is Evolving for Massive AI Workloads  | Ray Summit 2025RDMA P2P Deep Dive: KvCache Transfer, Weight Updates & MoE Routing at Perplexity | Ray Summit 2025Ray Direct Transport: RDMA Support in Ray Core | Ray Summit 2025Building LLaMA: Metas Director of GenAI Sergey Edunov | Ray Summit 2024Maximizing Compute Efficiency on Anyscale | Ray Summit 2025The PARK Stack: The LAMP Stack of the AI Era | Ben LoricaAgentic Workload Inference at Scale: ByteDance’s AIBrix & DeerFlow | Ray Summit 2025
Anyscale |

How Workday Achieved 50x Cheaper Model Serving with Ray Serve | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER