Squeezing Every Token Out of Your GPUs with llm-d @RedHatOpen
Squeezing Every Token Out of Your GPUs with llm-d  @RedHatOpen
Uploaded August 2026 | Updated September 2026, 19 hours ago
Large Language Models require incredibly expensive hardware to run. As AI workloads become more complex—especially with long-context tasks like agentic coding that generate thousands of tokens per session—keeping costs down is a major challenge.

In this video short, discover how llm-d implements key performance optimizations to help you squeeze maximum token efficiency out of your GPUs.

Additional Resources:
What is llm-d and why do we need it? redhat.com/en/blog/what-llm-d-and-why-do-we-need-it

What is llm-d: redhat.com/en/topics/ai/what-is-llm-d


The future of AI should be open. The Red Hat Open Source and AI Program Office (OSAIPO) builds, champions, and sustains Red Hat's open source leadership and engagement in the AI era. We guide communities in the responsible integration of AI to accelerate innovation, increase collaboration, and shape open standards.

#opensource #artificialintelligence #machinelearning
Squeezing Every Token Out of Your GPUs with llm-dFrom Docker Compose to Kubernetes with PodmanOCI Artifacts: Adding Support for Reference TypesChallenges of Using User Namespaces at Big ScaleValue of Open Source AILive Hardware Development at UCSC - Red Hat Research Days US 2020The Open Road: Does Onboarding Ever Stop?Using AI in open source: Quality over quantityAnalyzing the security certifications landscape: Does certification help security?Steps Toward Open Source Education - Red Hat Research Days 2021Red Hat NEXT! 2022 Keynote: The Future of AI Edge and Security is NowThe Open Road: DEI and Community Anonymity
Red Hat Open |

Squeezing Every Token Out of Your GPUs with llm-d

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER