Uploaded August 2026 | Updated September 2026, 2 weeks ago
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io
From Model Serving To Distributed Inference: How Llm-d Evolves AI Platforms on Kubernetes - Kay Yan, DaoCloud & Linbo He, Microsoft
This session explores how platform teams can evolve from basic model serving to production-grade distributed inference on Kubernetes. KAITO streamlines model onboarding and lifecycle management with curated presets, GPU node auto-provisioning, autoscaling, OpenAI-compatible endpoints, and support for runtimes such as vLLM. llm-d extends this foundation with inference-aware scheduling, KV-cache-aware routing, cross-node cache coordination, prefill/decode disaggregation, and workload-aware autoscaling. Rather than treating deployment and inference routing as separate concerns, this talk presents a practical reference architecture that brings both together on Kubernetes. It also explains how Gateway API Inference Extension can serve as a common interface, including InferencePool, Endpoint Picker, and Body Based Routing, and when teams should move from basic serving to distributed inference.
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io
From Model Serving To Distributed Inference: How Llm-d Evolves AI Platforms on Kubernetes - Kay Yan, DaoCloud & Linbo He, Microsoft
This session explores how platform teams can evolve from basic model serving to production-grade distributed inference on Kubernetes. KAITO streamlines model onboarding and lifecycle management with curated presets, GPU node auto-provisioning, autoscaling, OpenAI-compatible endpoints, and support for runtimes such as vLLM. llm-d extends this foundation with inference-aware scheduling, KV-cache-aware routing, cross-node cache coordination, prefill/decode disaggregation, and workload-aware autoscaling. Rather than treating deployment and inference routing as separate concerns, this talk presents a practical reference architecture that brings both together on Kubernetes. It also explains how Gateway API Inference Extension can serve as a common interface, including InferencePool, Endpoint Picker, and Body Based Routing, and when teams should move from basic serving to distributed inference.










