LiquidAI’s Approach to Large-Scale Synthetic Data Generation Using Ray | Ray Summit 2025 @anyscale
LiquidAI’s Approach to Large-Scale Synthetic Data Generation Using Ray | Ray Summit 2025  @anyscale
Uploaded December 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Arthur Book from Liquid.ai shares practical design patterns for combining Ray Data, Ray Serve, and vLLM to build scalable, high-throughput pipelines for synthetic data generation—an increasingly essential component of modern LLM development.

He begins by outlining the challenges of generating synthetic data at scale, where teams must coordinate large numbers of inference calls, manage multi-step agentic workflows, and maintain reliable throughput across heterogeneous GPU clusters. Arthur demonstrates how Ray’s unified execution model enables these workloads to run efficiently without complex, ad-hoc orchestration.

The session then dives into concrete implementation strategies, including:

Leveraging Ray Data for distributed data ingestion, transformation, batching, and parallelization

Using Ray Serve + vLLM for high-performance inference across multiple agents

Integrating agents into multi-step refinement loops, ensuring correctness and improving data quality

Managing GPU allocation, backpressure, and autoscaling in synthetic data pipelines

As a hands-on example, Arthur walks through building a two-agent self-refinement loop powered by Ray Serve and vLLM, and shows how it seamlessly integrates into a Ray Data workflow to create a robust, end-to-end synthetic data generation pipeline.

Attendees will walk away with actionable patterns for constructing scalable synthetic data systems—and a deeper understanding of how Ray’s components combine to power complex, high-throughput LLM workflows.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
LiquidAI’s Approach to Large-Scale Synthetic Data Generation Using Ray | Ray Summit 2025Why RAG Breaks at Scale | Anyscale WebinarRay Summit 2025 Keynote: Vehicle Intelligence at Scale with Peter Ludwig from Applied IntuitionSasha Rush on Building Cursor Composer and the Future of Agentic CodingScaling Multi-Modal Datasets to Petabytes with Ray at Apple | Ray Summit 2025Ray: Last Year’s Progress and the Road Ahead | Ray Summit 2025SGLang: An Efficient Open-Source Framework for Large-Scale LLM Serving | Ray Summit 2025CoServe: Max Performance, Minimal Compute | Ray Summit 2025How Ray Data Powers Scalable AI Workloads | Ray Summit 2025Inside Uber: Scaling Model Training with Ray | Ray Summit 2025Transforming Multimodal Data Management with LanceDB-Ray | Ray Summit 2024Why Ray Became a Distributed Computing Engine for Modern AI
Anyscale |

LiquidAI’s Approach to Large-Scale Synthetic Data Generation Using Ray | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER