How DigitalOcean Builds Next-Gen Inference with Ray, vLLM & More | Ray Summit 2025 @anyscale
How DigitalOcean Builds Next-Gen Inference with Ray, vLLM & More | Ray Summit 2025  @anyscale
Uploaded December 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Yogesh Sharma, Boopathy Kannappan, and Debarshi Raha from DigitalOcean share how they built a robust, scalable inference platform for next-generation generative models—powered by Ray and vLLM, running on Kubernetes, and optimized for both serverless and dedicated GPU workloads.

They begin by outlining the rising complexity of inference as models grow in size, context length, and modality. Meeting real-world performance and reliability requirements demands a platform that can scale elastically, manage GPU resources intelligently, and handle dynamic workloads efficiently.

The speakers introduce DigitalOcean’s inference architecture, showing how:

Ray’s scheduling primitives ensure reliable execution across distributed clusters
Placement groups guarantee GPU affinity and predictable performance
Ray observability tools enable deep insight into system health and workload behavior
vLLM provides fast token streaming, optimized batching, and advanced memory/KV-cache management
Serverless and Dedicated Inference Modes

They explore two key operational modes:

Serverless inference for automatic scaling, burst handling, and cost efficiency
Dedicated inference for fine-grained GPU partitioning, custom quantization pipelines, and performance isolation

This dual-mode architecture allows DigitalOcean to serve diverse customer workloads while maintaining reliability and performance under varying traffic patterns.

Advanced Optimization for Long-Context Models

The team then discusses their ongoing initiatives to improve inference for models with contexts exceeding 8k tokens, including:

Dynamic batching by token length
KV cache reuse strategies
Speculative decoding to improve latency and throughput without sacrificing accuracy
Roadmap: Multimodal, Multi-Tenant, and Unified Orchestration
Finally, they present their roadmap for a fully multimodal, multi-tenant inference platform, featuring:

Concurrent model orchestration

Tenant isolation and security-aware billing

A vision for a centralized orchestration layer with Ray as the control plane

A unified model registry for intelligent model placement, prioritization, and lifecycle management

This talk is designed for AI infrastructure engineers building scalable inference systems—whether you're optimizing cutting-edge production stacks or just beginning to architect your own.

Attendees will leave with a clear understanding of how to build future-ready inference platforms capable of serving large, dynamic, multimodal generative models at scale.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
How DigitalOcean Builds Next-Gen Inference with Ray, vLLM & More | Ray Summit 2025Scaling User-Focused Foundation Models at Grab with Ray | Ray Summit 2025Ray Meets Daft: Supercharging ETL and Analytics | Ray Summit 2024Reverbs ML Evolution: From Data Engineering to MLOps | Ray Summit 2024Best Practices for Ray in ProductionNVIDIA NeMo Curator: Scaling Multi-Modal Data Curation Workflows | Ray Summit 2025Multi-tenant Data Processing with Ray: Phaidras Approach to Industrial AI | Ray Summit 2024How Roblox Trains 3D Foundation Models with Ray | Ray Summit 2025How Daft Boosts Batch Inference Throughput with Dynamic Partitioning | Ray Summit 2025Optimizing vLLM Performance through Quantization | Ray Summit 2024How IBM Research Achieved vLLM Platform Portability with Triton Autotuning | Ray Summit 2024Greg Brockman on Founding OpenAI and Systems for AI | Ray Summit 2022
Anyscale |

How DigitalOcean Builds Next-Gen Inference with Ray, vLLM & More | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER