Uploaded November 2025 | Updated September 2026, 1 week ago
At Ray Summit 2025, Haocheng Bian, Yihao Guo, and Arun Ananthampalayam from Apple share how to build a unified, flexible ML framework on Ray—capable of orchestrating highly complex workloads ranging from large-scale batch inference to training hundreds-of-billions–parameter models.
They begin by outlining the rising need for an ML platform that can seamlessly manage diverse workloads while maintaining reliability, efficiency, and developer productivity. The speakers walk through the architecture of Apple’s Ray-based framework, designed to unify training, inference, and data processing under a single, scalable system.
The session dives into how the team addressed major challenges encountered at extreme scale, including:
Resource optimization across heterogeneous infrastructure
Fault tolerance and resiliency for long-running, high-stakes workloads
Low-latency execution even as workloads grow in complexity
Balancing flexibility with ease of use for research and production teams
They demonstrate how Ray’s distributed computing model unlocks both performance and adaptability, enabling Apple teams to operationalize massive ML systems without sacrificing rigor, reliability, or speed.
Attendees will walk away with practical patterns for building scalable ML frameworks on Ray—as well as insights on deploying, managing, and optimizing large, heterogeneous AI workloads in production.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
At Ray Summit 2025, Haocheng Bian, Yihao Guo, and Arun Ananthampalayam from Apple share how to build a unified, flexible ML framework on Ray—capable of orchestrating highly complex workloads ranging from large-scale batch inference to training hundreds-of-billions–parameter models.
They begin by outlining the rising need for an ML platform that can seamlessly manage diverse workloads while maintaining reliability, efficiency, and developer productivity. The speakers walk through the architecture of Apple’s Ray-based framework, designed to unify training, inference, and data processing under a single, scalable system.
The session dives into how the team addressed major challenges encountered at extreme scale, including:
Resource optimization across heterogeneous infrastructure
Fault tolerance and resiliency for long-running, high-stakes workloads
Low-latency execution even as workloads grow in complexity
Balancing flexibility with ease of use for research and production teams
They demonstrate how Ray’s distributed computing model unlocks both performance and adaptability, enabling Apple teams to operationalize massive ML systems without sacrificing rigor, reliability, or speed.
Attendees will walk away with practical patterns for building scalable ML frameworks on Ray—as well as insights on deploying, managing, and optimizing large, heterogeneous AI workloads in production.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com










