llm-d 101: Why Do We Need It? @RedHatOpen
llm-d 101: Why Do We Need It?  @RedHatOpen
Uploaded August 2026 | Updated September 2026, 19 minutes ago
Red Hat's Brian Stevens and Robert Shaw discuss the fundamental financial and computational hurdles organizations face when scaling artificial intelligence (AI). They explain that as enterprise AI applications generate millions of inference requests across multiple models, tools, and agents, the industry's focus is rapidly shifting from model training to running models efficiently.

The discussion covers:
--The Stateless Load Balancing Problem: Why traditional Kubernetes load balancers, designed for stateless applications, fail when applied to LLM workloads.

--The KV Cache Bottleneck: How LLMs rely on Key-Value (KV) cache—the model's short-term memory—to avoid repeating work, and how "blind" conventional routing loses this cache locality. This failure forces GPUs to redundantly recompute prompts, resulting in unnecessary hardware utilization and high operational costs.

--vLLM & llm-d Collaboration: How organizations can resolve this by pairing vLLM (which optimizes single-node engine efficiency) with llm-d (acting as the cluster-wide control plane to orchestrate multiple vLLM instances as a unified platform)

Additional Resources:
What is llm-d and why do we need it? redhat.com/en/blog/what-llm-d-and-why-do-we-need-it
What is llm-d: redhat.com/en/topics/ai/what-is-llm-d


The future of AI should be open. The Red Hat Open Source and AI Program Office (OSAIPO) builds, champions, and sustains Red Hat's open source leadership and engagement in the AI era. We guide communities in the responsible integration of AI to accelerate innovation, increase collaboration, and shape open standards.

#opensource #artificialintelligence #redhat
llm-d 101: Why Do We Need It?Red Hat NEXT! 2022: Enabling Next Generation Networking with eBPFWhy AI Infrastructure Must Be Built in the OpenGetting Started with Inference Using vLLMCommunity Central: Using Pulp 3 - Simpler and FasterCoreOS vs Fedora IoTCommunity Central: What does the Continuous Delivery Foundation do?Combining Kubernetes and vLLM to Deliver Scalable, Distributed Inference with llm-dCommunity Central: Join It or Leave It? Key considerations for evaluating communities
Red Hat Open |

llm-d 101: Why Do We Need It?

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER