[vLLM Office Hours #55] - Mooncake + vLLM/llm-d Deep Dive - August 6, 2026 @redhat
[vLLM Office Hours #55] - Mooncake + vLLM/llm-d Deep Dive - August 6, 2026  @redhat
Uploaded August 2026 | Updated September 2026, 2 weeks ago
Welcome to vLLM office hours! These bi-weekly sessions are your chance to stay current with the vLLM ecosystem, ask questions, and hear directly from contributors and power users.

This week's special topic: Mooncake + vLLM/llm-d Deep Dive.

vLLM project update from core maintainer Michael Goin, covering the v0.26 release plus recent model support, performance work, and quantization updates.

Hackathon spotlight: Timothy and David present TickyTalky, a winning project from the Austin Red Hat and NVIDIA vLLM hackathon, built on DeepSeek V4 Flash running on a DGX Spark.

Then the main event, with Ke Yang and Christian Angelo Velarde (Approaching.AI) and Greg Pereira (Red Hat AI). Mooncake began as the inference architecture behind Kimi and has grown into communication and storage infrastructure for large-scale LLM serving. We cover how it integrates with vLLM for PD disaggregation and cluster-wide KV cache reuse, and how llm-d wraps these pieces for distributed serving on Kubernetes.

Slides: docs.google.com/presentation/d/1JCwGvhFdTJt_jCFvQYTmMp2zgWBVJIoSnxciYDP0U0o

Want to join the discussion live on Google Meet? Get a calendar invite by filling out this form: https://red.ht/office-hours

Timestamps:
00:00 Intro
02:48 About vLLM
06:50 Recent Community blogs and model releases
14:10 vLLM v0.26 release update
19:43 Hackathon spotlight: TickyTalky
26:38 Mooncake architecture and ecosystem
41:08 Mooncake x vLLM integration and benchmarks
51:24 llm-d: distributed serving and tiered cache offloading
[vLLM Office Hours #55] - Mooncake + vLLM/llm-d Deep Dive - August 6, 2026GitOps Guide to the Galaxy (ep 108) | Declarative Networking w/IsovalentIn the Clouds (E52) | Elevating the Enterprise: Red Hat Summit 2026 Preview ft. Chuck DubuqueData Security And AIMonday highlights from Red Hat Summit 2026What’s New In Red Hat AI 3.4?The path from CentOS Linux to RHELAI Explained: Reduce GPU costs with LLM CompressorLearn about secure enclaves for AI inferenceMove VMs forward without holding your business back60% less time in model training saved $2.5MThe key to token economics
Red Hat |

[vLLM Office Hours #55] - Mooncake + vLLM/llm-d Deep Dive - August 6, 2026

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER