From Model Serving To Distributed Inference: How Llm-d Evolves AI Platforms on… K. Yan & L. He @cncf
From Model Serving To Distributed Inference: How Llm-d Evolves AI Platforms on… K. Yan & L. He  @cncf
Uploaded August 2026 | Updated September 2026, 2 weeks ago
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io

From Model Serving To Distributed Inference: How Llm-d Evolves AI Platforms on Kubernetes - Kay Yan, DaoCloud & Linbo He, Microsoft

This session explores how platform teams can evolve from basic model serving to production-grade distributed inference on Kubernetes. KAITO streamlines model onboarding and lifecycle management with curated presets, GPU node auto-provisioning, autoscaling, OpenAI-compatible endpoints, and support for runtimes such as vLLM. llm-d extends this foundation with inference-aware scheduling, KV-cache-aware routing, cross-node cache coordination, prefill/decode disaggregation, and workload-aware autoscaling. Rather than treating deployment and inference routing as separate concerns, this talk presents a practical reference architecture that brings both together on Kubernetes. It also explains how Gateway API Inference Extension can serve as a common interface, including InferencePool, Endpoint Picker, and Body Based Routing, and when teams should move from basic serving to distributed inference.
From Model Serving To Distributed Inference: How Llm-d Evolves AI Platforms on… K. Yan & L. HeKeycloak + Sigstore: Binding Human Identity To Artifact Signatures - Oshi Gupta & Sagar UtekarA Decade of Cilium Around the World - Liz Rice, Isovalent at Cisco & Hiroki Hanada, Cybozu, Inc.Argo CD Maintainers Panel - Dan Garfield, Joanna Wyganowska, Michael Crenshaw & Nitish KumarVulnerability Response for Large Open Source Projects - Jo Guerreiro & Charline Voinot, Grafana LabsMeet Marcus Noble, CNCF AmbassadorSimon Forster: Why Non-Code Contributions Matter | KubeCon JapanRunning OpenSearch at Scale in High-Traffic Gaming Systems - Siddharth Vijay, Baazi GamesWhy CNCF Ambassador Sharma Shivlal Keeps Coming BackOpenTelemetry Celebrates Graduation and the Next Era of Agentic... - Alolita Sharma & Ted YoungProject Lightning Talk: Youki: Whats New and Whats Next? - Yuta Nagai, CyberAgent Inc.Turning Platform Engineering Work Into Business Value Leadership Understands - D. Cook & S. Forster
CNCF [Cloud Native Computing Foundation] |

From Model Serving To Distributed Inference: How Llm-d Evolves AI Platforms on… K. Yan & L. He

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER