Uploaded December 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Seth Kimmel from Sutro shares how Sutro—an accelerated batch inference service used for synthetic data generation, evaluations, and large-scale unstructured data processing—pushes vLLM to its limits to deliver predictable, high-throughput offline inference at massive scale.
He begins by outlining Sutro’s workload patterns, which span from a few hundred tokens to tens of billions per job. For these large offline workloads, predictability—in cost, performance, and execution transparency—is absolutely critical. To meet these requirements, Sutro has built a deeply optimized, vLLM-powered inference engine designed specifically for large batch processing.
Seth then dives into how Sutro uses vLLM under the hood, covering:
Custom internal implementation layers built on top of vLLM
A performance profiler that measures and predicts system behavior in real time
Throughput estimation algorithms that inform batching, scheduling, and hardware allocation
Cost attribution instrumentation that provides precise, job-level visibility into resource usage
These systems allow Sutro to reliably scale vLLM across enormous workloads while maintaining strict SLAs for customers.
This talk is ideal for teams aiming to push vLLM further—those operating at large batch sizes, generating synthetic datasets, or building evaluation pipelines where cost predictability and throughput consistency are essential. Attendees will walk away with practical techniques for designing transparent, high-performance vLLM infrastructure at scale.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
At Ray Summit 2025, Seth Kimmel from Sutro shares how Sutro—an accelerated batch inference service used for synthetic data generation, evaluations, and large-scale unstructured data processing—pushes vLLM to its limits to deliver predictable, high-throughput offline inference at massive scale.
He begins by outlining Sutro’s workload patterns, which span from a few hundred tokens to tens of billions per job. For these large offline workloads, predictability—in cost, performance, and execution transparency—is absolutely critical. To meet these requirements, Sutro has built a deeply optimized, vLLM-powered inference engine designed specifically for large batch processing.
Seth then dives into how Sutro uses vLLM under the hood, covering:
Custom internal implementation layers built on top of vLLM
A performance profiler that measures and predicts system behavior in real time
Throughput estimation algorithms that inform batching, scheduling, and hardware allocation
Cost attribution instrumentation that provides precise, job-level visibility into resource usage
These systems allow Sutro to reliably scale vLLM across enormous workloads while maintaining strict SLAs for customers.
This talk is ideal for teams aiming to push vLLM further—those operating at large batch sizes, generating synthetic datasets, or building evaluation pipelines where cost predictability and throughput consistency are essential. Attendees will walk away with practical techniques for designing transparent, high-performance vLLM infrastructure at scale.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute










