Building Fault-Tolerant Massive Ray Clusters on Anyscale | Ray Summit 2025 @anyscale
Building Fault-Tolerant Massive Ray Clusters on Anyscale | Ray Summit 2025  @anyscale
Uploaded November 2025 | Updated September 2026, 3 weeks ago
Slides: drive.google.com/file/d/1rnSSECmBeigaZDOxnbsq8GSh2mfZep55/view?usp=sharing

At Ray Summit 2025, Dhyey Shah and Ibrahim Rabbani from Anyscale share how Ray is engineered to withstand the harsh realities of large-scale AI workloads—and what it takes to run reliably on clusters exceeding 10,000 nodes.

They begin by outlining the challenges that inevitably arise at extreme scale: network flakiness, spot preemptions, hardware failures, resource contention, and unpredictable infrastructure behavior. The speakers explain how Ray’s distributed runtime is designed to absorb these failures gracefully, keep workloads running, and maintain strong reliability even under constant churn.

Dhyey and Ibrahim then dive into key engineering lessons learned from building and operating Ray at massive scale, including strategies for fault tolerance, state management, recovery, elasticity, and workload-aware scheduling.

Finally, they share what’s next for Ray—highlighting upcoming improvements to scalability, reliability, and performance aimed at powering the next generation of AI applications.

Whether you're running large-scale model training, distributed inference, reinforcement learning, or multimodal pipelines, this session offers a deep look at how Ray stays resilient when everything else breaks.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
Building Fault-Tolerant Massive Ray Clusters on Anyscale | Ray Summit 2025Scaling LLM Inference: AWS Inferentia Meets Ray Serve on EKS | Ray Summit 2024Inside NVIDIA Dynamo: Faster, Scalable AI Deployment | Ray Summit 2025Ray + Kubernetes: The Distributed OS for AI/ML | Ray on the Road – NYC 2025Ray Summit 2025 Keynote: Physical AI Turing Test with Jim Fan from NVIDIABen Horowitz - Historical Perspectives on AI and the Internet | Ray Summit 2023Ray Joins The Linux Foundation & PyTorch Sub-Foundation: Toward a Unified AI Compute StackHow Torc Robotics Scales Multimodal AI for Autonomous Driving with RayRay on Kubernetes: Powering Quant Research at Scale | Ray Summit 2024How the VAST AI Operating System Powers a Dynamic Data Plane for Ray | Ray Summit 2025Scaling Ray Train to 10K Kubernetes Nodes on GKE | Ray Summit 2024Motional’s Blueprint for High-Performance ML Systems in Autonomous Driving | Ray Summit 2025
Anyscale |

Building Fault-Tolerant Massive Ray Clusters on Anyscale | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER