KubeRay + vLLM at DatalogyAI: Engineering Trillion-Scale Synthetic Data Systems | Ray Summit 2025 @anyscale
KubeRay + vLLM at DatalogyAI: Engineering Trillion-Scale Synthetic Data Systems | Ray Summit 2025  @anyscale
Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Bogdan Gaza and Fan Pan from DatologyAI share how they are scaling synthetic data generation into a production-grade platform capable of processing trillions of tokens—a critical capability as high-quality training data becomes one of the biggest bottlenecks in advancing modern language models.

They explain how synthetic data has become essential for creating diverse, targeted datasets that complement organic sources, powering everything from fast 4.5B-parameter models to frontier-scale systems like GPT-5. To meet these demands, the DatologyAI team built a large-scale platform using KubeRay and vLLM, dynamically orchestrating thousands of GPU workers across multimodal tasks including recaptioning, rephrasing, and domain-specific content generation.

In this talk, Bogdan and Fan dive into key engineering lessons, including:

Driving vLLM inference to near-peak GPU utilization

Designing fault-tolerant Ray actors for tensor-parallel sharding

Auto-scaling KubeRay clusters to match shifting workload patterns

Applying storage and scheduling strategies that deliver both high performance and significant cost efficiency

They highlight practical patterns for building resilient, scalable ML infrastructure, showing how clean abstractions between Ray’s distributed computing layer and vLLM’s inference engine enable rapid iteration on prompt engineering while maintaining production stability.

Attendees will learn how research prototypes evolve into trillion-token synthetic datasets that define the next generation of AI capabilities.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
KubeRay + vLLM at DatalogyAI: Engineering Trillion-Scale Synthetic Data Systems | Ray Summit 2025The emerging OpenSource AI Stack for modern AI workloads ⚡  #aiinfrastructure #aiopsHow DigitalOcean Builds Next-Gen Inference with Ray, vLLM & More | Ray Summit 2025Scaling User-Focused Foundation Models at Grab with Ray | Ray Summit 2025Ray Meets Daft: Supercharging ETL and Analytics | Ray Summit 2024Reverbs ML Evolution: From Data Engineering to MLOps | Ray Summit 2024Best Practices for Ray in ProductionNVIDIA NeMo Curator: Scaling Multi-Modal Data Curation Workflows | Ray Summit 2025Multi-tenant Data Processing with Ray: Phaidras Approach to Industrial AI | Ray Summit 2024How Roblox Trains 3D Foundation Models with Ray | Ray Summit 2025How Daft Boosts Batch Inference Throughput with Dynamic Partitioning | Ray Summit 2025Optimizing vLLM Performance through Quantization | Ray Summit 2024
Anyscale |

KubeRay + vLLM at DatalogyAI: Engineering Trillion-Scale Synthetic Data Systems | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER