Uploaded May 2026 | Updated September 2026, 2 weeks ago
The gap between AI enthusiasm and AI in production is where most enterprise initiatives stall. In this keynote, Red Hat's leadership, customers, and partners show a different path, grounded in real AI journeys from frontier models to open source optimization, and proven in production at some of the world's most demanding enterprises. Hear how Red Hat's unified platform — spanning infrastructure, sovereignty, and intelligence — gives organizations the foundation to move fast with AI without betting on the wrong technology, the wrong vendor, or the wrong moment. The future isn't one cloud or one model. It’s the right platform.
0:00 Intro: From Linux to the AI Era
1:20 Matt Hicks: The Reality of Enterprise IT
5:39 The three inflection points: Linux, Cloud, AI
8:39 Red Hat’s AI Journey: Building agentic systems
13:20 Optimizing token economics with open source
14:50 The Future of Work: Developers, Managers, & AI
20:00 Ashesh Badani: Solving the virtualization cost crisis
26:10 Motorola Solutions and their mission-critical AI journey
34:30 Navigating digital sovereignty & global regulations
36:41 New Announcements: Red Hat Sovereign Cloud
39:00 Eurocontrol and sovereign aviation
41:28 Core 42 and UAE sovereign AI
50:17 Defining AI sovereignty: Model, Data, & Outcome
51:55 Telenor and the sovereign AI Factory
54:50 Chris Wright: The velocity of Open Source AI
59:10 The token economics of AI agents
1:01:10 Red Hat AI Enterprise: Metal to agents
1:11:20 BNP Paribas with a $600M AI value
1:14:09 NVIDIA & Red Hat AI Factory
1:22:40 Verizon and autonomous networks
1:32:43 Conclusion: Owning Your Intelligence
What are your key takeaways from the keynote? Let us know in the comments! Like this video and subscribe for more Red Hat insights.
Follow all the news from #RHSummit in the newsroom: https://red.ht/3RK1rD6
#RedHat #EnterpriseAI #HybridCloud #Virtualization #Automation #Sovereignty
The gap between AI enthusiasm and AI in production is where most enterprise initiatives stall. In this keynote, Red Hat's leadership, customers, and partners show a different path, grounded in real AI journeys from frontier models to open source optimization, and proven in production at some of the world's most demanding enterprises. Hear how Red Hat's unified platform — spanning infrastructure, sovereignty, and intelligence — gives organizations the foundation to move fast with AI without betting on the wrong technology, the wrong vendor, or the wrong moment. The future isn't one cloud or one model. It’s the right platform.
0:00 Intro: From Linux to the AI Era
1:20 Matt Hicks: The Reality of Enterprise IT
5:39 The three inflection points: Linux, Cloud, AI
8:39 Red Hat’s AI Journey: Building agentic systems
13:20 Optimizing token economics with open source
14:50 The Future of Work: Developers, Managers, & AI
20:00 Ashesh Badani: Solving the virtualization cost crisis
26:10 Motorola Solutions and their mission-critical AI journey
34:30 Navigating digital sovereignty & global regulations
36:41 New Announcements: Red Hat Sovereign Cloud
39:00 Eurocontrol and sovereign aviation
41:28 Core 42 and UAE sovereign AI
50:17 Defining AI sovereignty: Model, Data, & Outcome
51:55 Telenor and the sovereign AI Factory
54:50 Chris Wright: The velocity of Open Source AI
59:10 The token economics of AI agents
1:01:10 Red Hat AI Enterprise: Metal to agents
1:11:20 BNP Paribas with a $600M AI value
1:14:09 NVIDIA & Red Hat AI Factory
1:22:40 Verizon and autonomous networks
1:32:43 Conclusion: Owning Your Intelligence
What are your key takeaways from the keynote? Let us know in the comments! Like this video and subscribe for more Red Hat insights.
Follow all the news from #RHSummit in the newsroom: https://red.ht/3RK1rD6
#RedHat #EnterpriseAI #HybridCloud #Virtualization #Automation #Sovereignty






![[vLLM Office Hours #52] - vLLM Semantic Router: Safer, Faster, Multi-Model Inference - June 25, 2026
Welcome to vLLM office hours! These bi-weekly sessions are your chance to stay current with the vLLM ecosystem, ask questions, and hear directly from contributors and power users.
This weeks special topics: vLLM Semantic Router: Intelligent Routing for Safer, Faster, Multi-Model Inference, plus a first look at the new vLLM course on DeepLearning.AI.
We cover the latest vLLM project updates, including the v0.23 release (DeepSeek V4 hardening with TRT-LLM-gen attention kernels and EPLB for MoE, Model Runner V2 expanding to more models, FlashInfer sampler integration, unified Gemma 4 multimodal support with native MTP, new model support for Step 3.7 Flash, Cosmos 3 Reasoner, Granite Speech Plus, and JetBrains Mellum, pipeline parallelism optimizations, AMD RDNA3 quantization kernels, multi-tier KV cache offloading to disk and remote storage, ongoing Rust frontend work, a unified reasoning/tool-call parser interface, and async EPLB by default), plus community highlights on Minimax M3, Poolside Laguna M.1, GLM-5.2, and Prime Intellects RL at 1T Scale blog.
We also share a first look at the new, free vLLM course on Andrew Ngs DeepLearning.AI platform, with developer advocate Cedric Clyburn.
Then we dive into vLLM Semantic Router with Christopher Nuland, Chief AI Architect at Red Hat AI: how signal-driven routing works across cost, latency, privacy, safety, and modality; using confidence scores and anchor tuning to improve routing accuracy; combining semantic routing with PII redaction and guardrails to protect sensitive data; pairing it with llm-d for prefix-aware caching; its new upstream integration into Agent Gateway; and a live home-lab demo.
Slides: https://docs.google.com/presentation/d/1hqDA046cxyMbV_x8UX472xK9x3WsIn8c9513kulec_Y/
Want to join the discussion live on Google Meet? Get a calendar invite by filling out this form: https://red.ht/office-hours
Timestamps:
00:00 vLLM intro
01:12 vLLM v0.23 project update
21:24 New vLLM course on DeepLearning.AI
26:03 vLLM Semantic Router deep dive, with live demo
54:03 Q&A [vLLM Office Hours #52] - vLLM Semantic Router: Safer, Faster, Multi-Model Inference - June 25, 2026](https://i.ytimg.com/vi/QoHlqjSkNoo/mqdefault.jpg)



![[vLLM Office Hours #55] - Mooncake + vLLM/llm-d Deep Dive - August 6, 2026
Welcome to vLLM office hours! These bi-weekly sessions are your chance to stay current with the vLLM ecosystem, ask questions, and hear directly from contributors and power users.
This weeks special topic: Mooncake + vLLM/llm-d Deep Dive.
vLLM project update from core maintainer Michael Goin, covering the v0.26 release plus recent model support, performance work, and quantization updates.
Hackathon spotlight: Timothy and David present TickyTalky, a winning project from the Austin Red Hat and NVIDIA vLLM hackathon, built on DeepSeek V4 Flash running on a DGX Spark.
Then the main event, with Ke Yang and Christian Angelo Velarde (Approaching.AI) and Greg Pereira (Red Hat AI). Mooncake began as the inference architecture behind Kimi and has grown into communication and storage infrastructure for large-scale LLM serving. We cover how it integrates with vLLM for PD disaggregation and cluster-wide KV cache reuse, and how llm-d wraps these pieces for distributed serving on Kubernetes.
Slides: https://docs.google.com/presentation/d/1JCwGvhFdTJt_jCFvQYTmMp2zgWBVJIoSnxciYDP0U0o
Want to join the discussion live on Google Meet? Get a calendar invite by filling out this form: https://red.ht/office-hours
Timestamps:
00:00 Intro
02:48 About vLLM
06:50 Recent Community blogs and model releases
14:10 vLLM v0.26 release update
19:43 Hackathon spotlight: TickyTalky
26:38 Mooncake architecture and ecosystem
41:08 Mooncake x vLLM integration and benchmarks
51:24 llm-d: distributed serving and tiered cache offloading [vLLM Office Hours #55] - Mooncake + vLLM/llm-d Deep Dive - August 6, 2026](https://i.ytimg.com/vi/RFyeBEy1AP8/mqdefault.jpg)