Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Bogdan Gaza and Fan Pan from DatologyAI share how they are scaling synthetic data generation into a production-grade platform capable of processing trillions of tokens—a critical capability as high-quality training data becomes one of the biggest bottlenecks in advancing modern language models.
They explain how synthetic data has become essential for creating diverse, targeted datasets that complement organic sources, powering everything from fast 4.5B-parameter models to frontier-scale systems like GPT-5. To meet these demands, the DatologyAI team built a large-scale platform using KubeRay and vLLM, dynamically orchestrating thousands of GPU workers across multimodal tasks including recaptioning, rephrasing, and domain-specific content generation.
In this talk, Bogdan and Fan dive into key engineering lessons, including:
Driving vLLM inference to near-peak GPU utilization
Designing fault-tolerant Ray actors for tensor-parallel sharding
Auto-scaling KubeRay clusters to match shifting workload patterns
Applying storage and scheduling strategies that deliver both high performance and significant cost efficiency
They highlight practical patterns for building resilient, scalable ML infrastructure, showing how clean abstractions between Ray’s distributed computing layer and vLLM’s inference engine enable rapid iteration on prompt engineering while maintaining production stability.
Attendees will learn how research prototypes evolve into trillion-token synthetic datasets that define the next generation of AI capabilities.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
At Ray Summit 2025, Bogdan Gaza and Fan Pan from DatologyAI share how they are scaling synthetic data generation into a production-grade platform capable of processing trillions of tokens—a critical capability as high-quality training data becomes one of the biggest bottlenecks in advancing modern language models.
They explain how synthetic data has become essential for creating diverse, targeted datasets that complement organic sources, powering everything from fast 4.5B-parameter models to frontier-scale systems like GPT-5. To meet these demands, the DatologyAI team built a large-scale platform using KubeRay and vLLM, dynamically orchestrating thousands of GPU workers across multimodal tasks including recaptioning, rephrasing, and domain-specific content generation.
In this talk, Bogdan and Fan dive into key engineering lessons, including:
Driving vLLM inference to near-peak GPU utilization
Designing fault-tolerant Ray actors for tensor-parallel sharding
Auto-scaling KubeRay clusters to match shifting workload patterns
Applying storage and scheduling strategies that deliver both high performance and significant cost efficiency
They highlight practical patterns for building resilient, scalable ML infrastructure, showing how clean abstractions between Ray’s distributed computing layer and vLLM’s inference engine enable rapid iteration on prompt engineering while maintaining production stability.
Attendees will learn how research prototypes evolve into trillion-token synthetic datasets that define the next generation of AI capabilities.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com










