Scaling Ray Train to 10K Kubernetes Nodes on GKE | Ray Summit 2024 @anyscale
Scaling Ray Train to 10K Kubernetes Nodes on GKE | Ray Summit 2024  @anyscale
Uploaded October 2024 | Updated September 2026, 2 weeks ago
In the race to train larger and more complex AI models, the ability to scale efficiently across massive compute clusters is paramount. This session unveils Google's groundbreaking achievement in scaling Ray Train and Ray Data to an unprecedented 10,000 node Kubernetes cluster on Google Kubernetes Engine (GKE).

Andrew Sy Kim and Saikat Roychowdhury will take you on a deep dive into the architecture and innovations that made this feat possible. They'll explore the synergy between KubeRay and key Kubernetes enhancements, revealing how these tools work in concert to manage enormous distributed workloads. The talk will also spotlight advancements in GKE's GCS Fuse CSI driver, demonstrating its crucial role in enabling distributed checkpointing and dataset processing at scale. Attendees will gain valuable insights into pushing the boundaries of distributed machine learning infrastructure, applicable to both modest and massive deployments.

--

Interested in more?
- Watch the full Day 1 Keynote: youtu.be/jwZHJthQvXo
- Watch the full Day 2 Keynote youtu.be/Lury2ad6KG8

--

🔗 Connect with us:
- Subscribe to our YouTube channel: youtube.com/@anyscale
- Twitter: https://x.com/anyscalecompute
- LinkedIn: linkedin.com/company/joinanyscale
- Website: anyscale.com
Scaling Ray Train to 10K Kubernetes Nodes on GKE | Ray Summit 2024Motional’s Blueprint for High-Performance ML Systems in Autonomous Driving | Ray Summit 2025How Roblox Scaled Machine Learning by Leveraging Ray for Efficient Batch Inference | Ray Summit 2024Ray Train: Distributed Solutions for Removing Training Bottlenecks | Ray Summit 2025Pricing and Packaging Your AI Products for Scale | Ray Summit 2024Building a Multimodal Video Processing Pipeline with RayHow The AI Institute is Revolutionizing Robotics ML Training | Ray Summit 2024Why context engineering is going to play a big role in AI in the future #aiinfrastructureKubeRay + vLLM at DatalogyAI: Engineering Trillion-Scale Synthetic Data Systems | Ray Summit 2025The emerging OpenSource AI Stack for modern AI workloads ⚡  #aiinfrastructure #aiopsHow DigitalOcean Builds Next-Gen Inference with Ray, vLLM & More | Ray Summit 2025Scaling User-Focused Foundation Models at Grab with Ray | Ray Summit 2025
Anyscale |

Scaling Ray Train to 10K Kubernetes Nodes on GKE | Ray Summit 2024

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER