How Daft Boosts Batch Inference Throughput with Dynamic Partitioning | Ray Summit 2025 @anyscale
How Daft Boosts Batch Inference Throughput with Dynamic Partitioning | Ray Summit 2025  @anyscale
Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Kevin Wang from Eventual shares how Daft enables petabyte-scale multimodal query processing on Ray—unlocking high-performance batch inference as part of complex, end-to-end pipelines.

He begins by outlining a fundamental challenge in large-scale LLM inference: maximizing prefix caching without stalling GPUs. Traditional methods rely on a full upfront sort and partition of prompts to boost cache locality—but this preprocessing step leaves GPUs idle and limits overall throughput.

Daft solves this with a breakthrough technique called dynamic prefix partitioning. Instead of static pre-sorting, Daft continuously and automatically adjusts partitions in-flight as data streams into vLLM. This approach ensures:

High prefix cache hit rates without manual preprocessing

Full GPU saturation throughout the entire query

End-to-end performance gains across multimodal batch workloads

Kevin walks through how Daft integrates vLLM’s high-performance inference engine into its distributed execution model, powered by Ray. The session explores the inner workings of the Daft query optimizer and execution engine, along with performance benchmarks on real multimodal pipelines.

Attendees will learn how Daft’s architecture can transform large-scale batch inference—making multimodal data processing faster, more efficient, and dramatically easier to operate.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
How Daft Boosts Batch Inference Throughput with Dynamic Partitioning | Ray Summit 2025Optimizing vLLM Performance through Quantization | Ray Summit 2024How IBM Research Achieved vLLM Platform Portability with Triton Autotuning | Ray Summit 2024Greg Brockman on Founding OpenAI and Systems for AI | Ray Summit 2022Wisedocs’ Journey: Rebuilding & Accelerating ML with KubeRay | Ray Summit 2025[Ray Meetup] Ray + vLLM in Action: Lessons from Pinterest and Large Scale Distributed InferenceHow Zoox Built a Reliable, High-Velocity Model Serving Platform with Ray Serve | Ray Summit 2025Scaling Machine Learning at Tripadvisor: Our Journey with Ray and Anyscale | Ray Summit 2025AWS + vLLM: Building the Future of Open, Fast LLM Serving | Ray Summit 2025Ray + vLLM  Efficient Multi Node Orchestration for Sparse MoE Model Serving | Ray Summit 2025Hybrid RL + Imitation Learning for Robotics with Ray at RAI InstituteHow Runhouse Orchestrates Multi-Cluster Ray Workloads | Ray Summit 2025
Anyscale |

How Daft Boosts Batch Inference Throughput with Dynamic Partitioning | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER