Inside NVIDIA Dynamo: Faster, Scalable AI Deployment | Ray Summit 2025 @anyscale
Inside NVIDIA Dynamo: Faster, Scalable AI Deployment | Ray Summit 2025  @anyscale
Uploaded November 2025 | Updated September 2026, 3 weeks ago
At Ray Summit 2025, Harry Kim from NVIDIA shares how NVIDIA Dynamo is redefining large-scale LLM inference through system-level optimizations that seamlessly integrate with high-performance engines such as vLLM, SGLang, and TensorRT-LLM (TRT-LLM).

He begins by outlining the core challenge: as LLMs grow in size, context length, and real-world usage, inference systems must deliver massive efficiency gains—not just from kernels or hardware, but across the entire distributed serving stack. NVIDIA Dynamo addresses this by introducing a new layer of intelligent orchestration and memory management designed specifically for LLM workloads.

Harry walks through Dynamo’s key innovations, including:

Smart Scheduling – Routes requests based on KV-cache hit rates and system load, intelligently autoscaling and disaggregating the prefill and decode phases for maximum throughput and efficiency.

Hierarchical Memory Management – Transparently leverages HBM, CPU memory, local NVMe, and remote storage to minimize latency and maximize effective model capacity.

Low-Latency KV-Cache Transfer – Quickly moves KV-cache across nodes and memory tiers, enabling fast context reuse and efficient distributed inference.

The session also introduces Dynamo’s production-grade LLM serving capabilities, including:

Tools to identify optimal disaggregated serving configurations offline

Automated tuning based on real-time traffic

Topology-aware gang scheduling to dynamically scale prefill and decode workers

LLM-specific fault-tolerance mechanisms for reliable serving at scale

Harry demonstrates how Dynamo enables organizations to achieve higher throughput, lower latency, and better cost efficiency across distributed LLM deployments—while still leveraging their preferred inference engine.

Attendees will leave with a clear understanding of how NVIDIA Dynamo transforms end-to-end LLM serving, making large-scale inference more efficient, robust, and operationally simple.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
Inside NVIDIA Dynamo: Faster, Scalable AI Deployment | Ray Summit 2025Ray + Kubernetes: The Distributed OS for AI/ML | Ray on the Road – NYC 2025Ray Summit 2025 Keynote: Physical AI Turing Test with Jim Fan from NVIDIABen Horowitz - Historical Perspectives on AI and the Internet | Ray Summit 2023Ray Joins The Linux Foundation & PyTorch Sub-Foundation: Toward a Unified AI Compute StackHow Torc Robotics Scales Multimodal AI for Autonomous Driving with RayRay on Kubernetes: Powering Quant Research at Scale | Ray Summit 2024How the VAST AI Operating System Powers a Dynamic Data Plane for Ray | Ray Summit 2025Scaling Ray Train to 10K Kubernetes Nodes on GKE | Ray Summit 2024Motional’s Blueprint for High-Performance ML Systems in Autonomous Driving | Ray Summit 2025How Roblox Scaled Machine Learning by Leveraging Ray for Efficient Batch Inference | Ray Summit 2024Ray Train: Distributed Solutions for Removing Training Bottlenecks | Ray Summit 2025
Anyscale |

Inside NVIDIA Dynamo: Faster, Scalable AI Deployment | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER