Uploaded December 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Arthur Book from Liquid.ai shares practical design patterns for combining Ray Data, Ray Serve, and vLLM to build scalable, high-throughput pipelines for synthetic data generation—an increasingly essential component of modern LLM development.
He begins by outlining the challenges of generating synthetic data at scale, where teams must coordinate large numbers of inference calls, manage multi-step agentic workflows, and maintain reliable throughput across heterogeneous GPU clusters. Arthur demonstrates how Ray’s unified execution model enables these workloads to run efficiently without complex, ad-hoc orchestration.
The session then dives into concrete implementation strategies, including:
Leveraging Ray Data for distributed data ingestion, transformation, batching, and parallelization
Using Ray Serve + vLLM for high-performance inference across multiple agents
Integrating agents into multi-step refinement loops, ensuring correctness and improving data quality
Managing GPU allocation, backpressure, and autoscaling in synthetic data pipelines
As a hands-on example, Arthur walks through building a two-agent self-refinement loop powered by Ray Serve and vLLM, and shows how it seamlessly integrates into a Ray Data workflow to create a robust, end-to-end synthetic data generation pipeline.
Attendees will walk away with actionable patterns for constructing scalable synthetic data systems—and a deeper understanding of how Ray’s components combine to power complex, high-throughput LLM workflows.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
At Ray Summit 2025, Arthur Book from Liquid.ai shares practical design patterns for combining Ray Data, Ray Serve, and vLLM to build scalable, high-throughput pipelines for synthetic data generation—an increasingly essential component of modern LLM development.
He begins by outlining the challenges of generating synthetic data at scale, where teams must coordinate large numbers of inference calls, manage multi-step agentic workflows, and maintain reliable throughput across heterogeneous GPU clusters. Arthur demonstrates how Ray’s unified execution model enables these workloads to run efficiently without complex, ad-hoc orchestration.
The session then dives into concrete implementation strategies, including:
Leveraging Ray Data for distributed data ingestion, transformation, batching, and parallelization
Using Ray Serve + vLLM for high-performance inference across multiple agents
Integrating agents into multi-step refinement loops, ensuring correctness and improving data quality
Managing GPU allocation, backpressure, and autoscaling in synthetic data pipelines
As a hands-on example, Arthur walks through building a two-agent self-refinement loop powered by Ray Serve and vLLM, and shows how it seamlessly integrates into a Ray Data workflow to create a robust, end-to-end synthetic data generation pipeline.
Attendees will walk away with actionable patterns for constructing scalable synthetic data systems—and a deeper understanding of how Ray’s components combine to power complex, high-throughput LLM workflows.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale










