Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Kevin Wang from Eventual shares how Daft enables petabyte-scale multimodal query processing on Ray—unlocking high-performance batch inference as part of complex, end-to-end pipelines.
He begins by outlining a fundamental challenge in large-scale LLM inference: maximizing prefix caching without stalling GPUs. Traditional methods rely on a full upfront sort and partition of prompts to boost cache locality—but this preprocessing step leaves GPUs idle and limits overall throughput.
Daft solves this with a breakthrough technique called dynamic prefix partitioning. Instead of static pre-sorting, Daft continuously and automatically adjusts partitions in-flight as data streams into vLLM. This approach ensures:
High prefix cache hit rates without manual preprocessing
Full GPU saturation throughout the entire query
End-to-end performance gains across multimodal batch workloads
Kevin walks through how Daft integrates vLLM’s high-performance inference engine into its distributed execution model, powered by Ray. The session explores the inner workings of the Daft query optimizer and execution engine, along with performance benchmarks on real multimodal pipelines.
Attendees will learn how Daft’s architecture can transform large-scale batch inference—making multimodal data processing faster, more efficient, and dramatically easier to operate.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
At Ray Summit 2025, Kevin Wang from Eventual shares how Daft enables petabyte-scale multimodal query processing on Ray—unlocking high-performance batch inference as part of complex, end-to-end pipelines.
He begins by outlining a fundamental challenge in large-scale LLM inference: maximizing prefix caching without stalling GPUs. Traditional methods rely on a full upfront sort and partition of prompts to boost cache locality—but this preprocessing step leaves GPUs idle and limits overall throughput.
Daft solves this with a breakthrough technique called dynamic prefix partitioning. Instead of static pre-sorting, Daft continuously and automatically adjusts partitions in-flight as data streams into vLLM. This approach ensures:
High prefix cache hit rates without manual preprocessing
Full GPU saturation throughout the entire query
End-to-end performance gains across multimodal batch workloads
Kevin walks through how Daft integrates vLLM’s high-performance inference engine into its distributed execution model, powered by Ray. The session explores the inner workings of the Daft query optimizer and execution engine, along with performance benchmarks on real multimodal pipelines.
Attendees will learn how Daft’s architecture can transform large-scale batch inference—making multimodal data processing faster, more efficient, and dramatically easier to operate.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com




![[Ray Meetup] Ray + vLLM in Action: Lessons from Pinterest and Large Scale Distributed Inference
Listen in to our Ray Meetup where we explored batch inference at scale with Ray and vLLM! Learn how Pinterest scales batch inference using Ray, and get a first look at Anyscale’s latest tools—Ray Serve and Data LLM—for orchestrating large-scale LLM inference. We’ll cover topics like batch inference, prefill-decode disaggregation, DP/EP parallelism, and custom request routing.
Speakers:
Chia-Wei Chen, Software Engineer, ML Training Infra, Pinterest
Kourosh Hakhamaneshi, AI Lead, Anyscale [Ray Meetup] Ray + vLLM in Action: Lessons from Pinterest and Large Scale Distributed Inference](https://i.ytimg.com/vi/HDSy09hrm2I/mqdefault.jpg)





