Scaling Production LLM Inference Using EKS Auto Mode & Ray Serve | Ray Summit 2025 @anyscale
Scaling Production LLM Inference Using EKS Auto Mode & Ray Serve | Ray Summit 2025  @anyscale
Uploaded December 2025 | Updated September 2026, 1 week ago
At Ray Summit 2025, Apoorva Kulkarni from AWS shares how teams can run large-scale LLM inference without becoming Kubernetes experts—by combining EKS Auto Mode with Ray Serve to create a fully automated, production-grade serving platform.

He begins by highlighting a common reality: teams often get stuck managing clusters, tuning autoscalers, wrangling GPU nodes, and troubleshooting infrastructure instead of focusing on building AI applications. EKS Auto Mode changes this dynamic by eliminating nearly all operational overhead.

Apoorva walks through a real-world deployment transformation—moving from labor-intensive, manually managed clusters to a self-healing, cost-efficient, and fully automated LLM serving platform. The session demonstrates how EKS Auto Mode provides:

Intelligent node provisioning tailored to AI workloads

Automatic, workload-driven scaling for both CPUs and GPUs

Built-in observability for system and workload visibility

Seamless GPU lifecycle management for inference-heavy pipelines

Burst capacity handling to maintain low latency under unpredictable load

Cost optimization for expensive inference accelerators

With Ray Serve orchestrating high-throughput, multi-model LLM inference on top, the result is a resilient system that scales from prototype to production with minimal complexity.

Attendees will leave with a clear blueprint for deploying scalable, reliable, and cost-efficient LLM inference on AWS—ideal for ML engineers who want to stop wrestling with Kubernetes and platform teams seeking turnkey AI infrastructure solutions.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
Scaling Production LLM Inference Using EKS Auto Mode & Ray Serve | Ray Summit 2025How Autodesk Built a Next-Gen Deep Learning Platform with Ray | Ray Summit 2025Building Scalable AI Infrastructure with Kuberay and Kubernetes | Ray Summit 2024Scaling AI at Autodesk with Ray and Metaflow | Ray Summit 2024Pinterest’s Approach to Real-Time ML Experimentation Using Ray | Ray Summit 2025Ray Summit 2025 Keynote: AI OSS Stack Panel with vLLM + PyTorch + KubernetesThe Future of AI Infrastructure: Anyscale Keynote | Ray on the Road – NYC 2025Ray Data for Structured Workloads: Deep Dive | Ray Summit 2025Ray Summit 2025: Jimmy Ba on Efficient Teams, AI Research, and What’s NextScaling LinkedIns Online Training Solution with Ray | Ray Summit 2025AI workloads introduce a system design shift ...Scaling LLMs on Google Cloud: Synergy Between Ray, TPU, and GKE | Ray Summit 2024
Anyscale |

Scaling Production LLM Inference Using EKS Auto Mode & Ray Serve | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER