Beyond VLLM: Distributed LLM Inferencing With Llm-d on Kubernetes - Ravindra Patil, Red Hat @cncf
Beyond VLLM: Distributed LLM Inferencing With Llm-d on Kubernetes - Ravindra Patil, Red Hat  @cncf
Uploaded July 2026 | Updated September 2026, 3 weeks ago
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Yokohama, Japan (29-30 July, 2026), and Shanghai, China (8-9 September, 2026) Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io

Beyond VLLM: Distributed LLM Inferencing With Llm-d on Kubernetes - Ravindra Patil, Red Hat

As (LLMs) continue to grow in size and demand, single-node inferencing quickly becomes a bottleneck for performance, scalability, and cost. While vLLM has become popular for efficient LLM serving on a single node, it does not fully address the challenges of distributed inferencing across multiple GPUs and nodes in Kubernetes environments.

This talk introduces llm-d, a emerging cloud-native project designed to enable distributed LLM inferencing on Kubernetes. We will cover why vLLM gained popularity and the limitations when scaling beyond a single node. We will explore how llm-d goes a step further by enabling multi-node, multi-GPU inferencing with cloud-native primitives.

Attendees will learn how llm-d fits into modern Kubernetes platforms, how it improves scalability and resource utilization. The session focuses on practical architecture, design trade-offs, and real-world use cases rather than theory with a demo on how llm-d distributes load.
Beyond VLLM: Distributed LLM Inferencing With Llm-d on Kubernetes - Ravindra Patil, Red HatSBOMit: Making SBOMs Accurate With Attestations - Marco De Vincenzi & Justin Cappos, NYUProject Lightning Talk: Opening Remarks - Hoon Jo, KubernetesLab.devBeyond Single-Cluster Limits: Scaling GPU Workloads Across Kubernetes With… K. Das & E. BayramovaDetecting Compromised CI With eBPF and Cilium Tetragon - Liz Rice, Isovalent at CiscoShared Yet Isolated at Scale: Building Multi-Tenant Inference Platform on… Y. Hiraki & Y. TanakaNeurodiversity Meeting - December 2025The Death of the YAML-Engineer: Engineering Invisible Platf... Abhinav Sharma & Mumshad Mannambeth⚡Lightning Talk: Learning Kubernetes From Logs: Building a Foundation for Future… K. IshiiWhat Is a Virtual Power Plant (VPP) ? : Green Tech and the Modernization of the Grid - L. MalcomSIG API Machinery in the Era of AI: Updates - Federico Bongiovanni, Google CloudRunning Wasm Inside Your Storage Cluster With CSI and Gateway API - Ho Kim & SangJoon Park
CNCF [Cloud Native Computing Foundation] |

Beyond VLLM: Distributed LLM Inferencing With Llm-d on Kubernetes - Ravindra Patil, Red Hat

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER