Uploaded November 2025 | Updated September 2026, 1 week ago
At Ray Summit 2025, Bharat Joshi and Peng Zhang from Uber share how Ray has become a cornerstone of Uber’s rapidly expanding machine learning infrastructure—powering scalable, flexible, and efficient distributed training across heterogeneous compute environments.
They begin by outlining Uber’s growing need to support large-scale training for LLMs, recommendation systems, and other high-capacity models, where traditional systems struggled to meet evolving performance and reliability requirements. Ray’s unified distributed computing model has enabled Uber to standardize training workflows while operating at massive scale.
The speakers then dive into the architectural strategies that have been critical to Uber’s success, including:
Multi-cloud training, allowing workloads to span diverse compute providers for improved flexibility and resource availability
Disaster recovery–ready design, ensuring continuity for production-grade ML workloads even in the face of cloud outages or regional failures
Cross-environment portability, supporting seamless transitions between cloud and on-premise clusters
They also share detailed learnings from optimizing performance and training throughput—covering GPU/accelerator efficiency techniques, system-level tuning, and the patterns that proved most effective for large distributed jobs.
Attendees will gain practical insights into how Uber leverages Ray to train large-scale models reliably across complex, mixed compute environments—and how these architectural lessons can be applied to enterprise ML systems of any scale.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
At Ray Summit 2025, Bharat Joshi and Peng Zhang from Uber share how Ray has become a cornerstone of Uber’s rapidly expanding machine learning infrastructure—powering scalable, flexible, and efficient distributed training across heterogeneous compute environments.
They begin by outlining Uber’s growing need to support large-scale training for LLMs, recommendation systems, and other high-capacity models, where traditional systems struggled to meet evolving performance and reliability requirements. Ray’s unified distributed computing model has enabled Uber to standardize training workflows while operating at massive scale.
The speakers then dive into the architectural strategies that have been critical to Uber’s success, including:
Multi-cloud training, allowing workloads to span diverse compute providers for improved flexibility and resource availability
Disaster recovery–ready design, ensuring continuity for production-grade ML workloads even in the face of cloud outages or regional failures
Cross-environment portability, supporting seamless transitions between cloud and on-premise clusters
They also share detailed learnings from optimizing performance and training throughput—covering GPU/accelerator efficiency techniques, system-level tuning, and the patterns that proved most effective for large distributed jobs.
Attendees will gain practical insights into how Uber leverages Ray to train large-scale models reliably across complex, mixed compute environments—and how these architectural lessons can be applied to enterprise ML systems of any scale.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com









