Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Steve Han, Yiqing Wang, and Liangjun Feng from Roblox share how the company built a modern ML platform on Ray and used it to train their 3D foundation model at scale.
They begin by walking through the platform architecture, including how Roblox integrated KubeRay with Istio and Kubeflow to support authentication, multi-tenancy, and secure workflow orchestration. They also discuss open-sourcing the new KubeRay dashboard, designed to improve iterative development, along with innovations such as p2p Docker image distribution, lazy image pulling, and scaling Ray jobs across multiple Kubernetes clusters.
The speakers then explore the challenges of applying Ray to large-scale foundation model training. This includes running high-volume LLM batch labeling jobs, leveraging Ray Data at scale, and supporting demanding distributed workloads. Historically, Roblox launched distributed training with MPI, which provided the basics of multi-GPU execution but lacked critical features like observability and fault tolerance. Over the past year, the Roblox platform team evaluated and adopted Ray Train as their default distributed training framework—successfully migrating the majority of training use cases to Ray.
Attendees will gain insights into building a production ML platform on Ray, modernizing distributed training infrastructure, and supporting large-scale foundation model development in a complex, multi-team environment.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
đź”— Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
At Ray Summit 2025, Steve Han, Yiqing Wang, and Liangjun Feng from Roblox share how the company built a modern ML platform on Ray and used it to train their 3D foundation model at scale.
They begin by walking through the platform architecture, including how Roblox integrated KubeRay with Istio and Kubeflow to support authentication, multi-tenancy, and secure workflow orchestration. They also discuss open-sourcing the new KubeRay dashboard, designed to improve iterative development, along with innovations such as p2p Docker image distribution, lazy image pulling, and scaling Ray jobs across multiple Kubernetes clusters.
The speakers then explore the challenges of applying Ray to large-scale foundation model training. This includes running high-volume LLM batch labeling jobs, leveraging Ray Data at scale, and supporting demanding distributed workloads. Historically, Roblox launched distributed training with MPI, which provided the basics of multi-GPU execution but lacked critical features like observability and fault tolerance. Over the past year, the Roblox platform team evaluated and adopted Ray Train as their default distributed training framework—successfully migrating the majority of training use cases to Ray.
Attendees will gain insights into building a production ML platform on Ray, modernizing distributed training infrastructure, and supporting large-scale foundation model development in a complex, multi-team environment.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
đź”— Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com





![[Ray Meetup] Ray + vLLM in Action: Lessons from Pinterest and Large Scale Distributed Inference
Listen in to our Ray Meetup where we explored batch inference at scale with Ray and vLLM! Learn how Pinterest scales batch inference using Ray, and get a first look at Anyscale’s latest tools—Ray Serve and Data LLM—for orchestrating large-scale LLM inference. We’ll cover topics like batch inference, prefill-decode disaggregation, DP/EP parallelism, and custom request routing.
Speakers:
Chia-Wei Chen, Software Engineer, ML Training Infra, Pinterest
Kourosh Hakhamaneshi, AI Lead, Anyscale [Ray Meetup] Ray + vLLM in Action: Lessons from Pinterest and Large Scale Distributed Inference](https://i.ytimg.com/vi/HDSy09hrm2I/mqdefault.jpg)




