Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Jacob Huffman and Hao Wang from NVIDIA share how Roblox built a modern ML platform on Ray and leveraged it to train their large-scale 3D foundation model.
They begin by walking through the platform’s architecture, including how Roblox integrated KubeRay with Istio and Kubeflow to support authentication, multi-tenancy, and secure orchestration. They also highlight efforts to open-source the new KubeRay dashboard, designed to improve iterative development, along with infrastructure innovations such as p2p Docker image distribution, lazy pulling, and the ability to scale Ray jobs across multiple clusters.
Jacob and Hao then dive into the challenges Roblox faced when applying Ray to foundation-model training workloads. This includes orchestrating massive LLM batch labeling jobs, leveraging Ray Data at scale, and supporting large distributed pipelines across heterogeneous compute. Historically, Roblox relied on MPI to launch distributed training, which handled multi-GPU execution but lacked critical capabilities like observability and fault tolerance.
Over the past year, the Roblox platform team evaluated and adopted Ray Train as their default distributed training framework—successfully migrating the majority of their training workloads from MPI to Ray, improving reliability and simplifying operations.
Attendees will take away practical insights into building production-grade Ray platforms, modernizing distributed training workflows, and supporting multimodal foundation model development at scale.
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
At Ray Summit 2025, Jacob Huffman and Hao Wang from NVIDIA share how Roblox built a modern ML platform on Ray and leveraged it to train their large-scale 3D foundation model.
They begin by walking through the platform’s architecture, including how Roblox integrated KubeRay with Istio and Kubeflow to support authentication, multi-tenancy, and secure orchestration. They also highlight efforts to open-source the new KubeRay dashboard, designed to improve iterative development, along with infrastructure innovations such as p2p Docker image distribution, lazy pulling, and the ability to scale Ray jobs across multiple clusters.
Jacob and Hao then dive into the challenges Roblox faced when applying Ray to foundation-model training workloads. This includes orchestrating massive LLM batch labeling jobs, leveraging Ray Data at scale, and supporting large distributed pipelines across heterogeneous compute. Historically, Roblox relied on MPI to launch distributed training, which handled multi-GPU execution but lacked critical capabilities like observability and fault tolerance.
Over the past year, the Roblox platform team evaluated and adopted Ray Train as their default distributed training framework—successfully migrating the majority of their training workloads from MPI to Ray, improving reliability and simplifying operations.
Attendees will take away practical insights into building production-grade Ray platforms, modernizing distributed training workflows, and supporting multimodal foundation model development at scale.
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com










