[vLLM Office Hours #52] - vLLM Semantic Router: Safer, Faster, Multi-Model Inference - June 25, 2026 @redhat
[vLLM Office Hours #52] - vLLM Semantic Router: Safer, Faster, Multi-Model Inference - June 25, 2026  @redhat
Uploaded July 2026 | Updated September 2026, 2 weeks ago
Welcome to vLLM office hours! These bi-weekly sessions are your chance to stay current with the vLLM ecosystem, ask questions, and hear directly from contributors and power users.

This week's special topics: vLLM Semantic Router: Intelligent Routing for Safer, Faster, Multi-Model Inference, plus a first look at the new vLLM course on DeepLearning.AI.

We cover the latest vLLM project updates, including the v0.23 release (DeepSeek V4 hardening with TRT-LLM-gen attention kernels and EPLB for MoE, Model Runner V2 expanding to more models, FlashInfer sampler integration, unified Gemma 4 multimodal support with native MTP, new model support for Step 3.7 Flash, Cosmos 3 Reasoner, Granite Speech Plus, and JetBrains Mellum, pipeline parallelism optimizations, AMD RDNA3 quantization kernels, multi-tier KV cache offloading to disk and remote storage, ongoing Rust frontend work, a unified reasoning/tool-call parser interface, and async EPLB by default), plus community highlights on Minimax M3, Poolside Laguna M.1, GLM-5.2, and Prime Intellect's "RL at 1T Scale" blog.

We also share a first look at the new, free vLLM course on Andrew Ng's DeepLearning.AI platform, with developer advocate Cedric Clyburn.
Then we dive into vLLM Semantic Router with Christopher Nuland, Chief AI Architect at Red Hat AI: how signal-driven routing works across cost, latency, privacy, safety, and modality; using confidence scores and anchor tuning to improve routing accuracy; combining semantic routing with PII redaction and guardrails to protect sensitive data; pairing it with llm-d for prefix-aware caching; its new upstream integration into Agent Gateway; and a live home-lab demo.

Slides: docs.google.com/presentation/d/1hqDA046cxyMbV_x8UX472xK9x3WsIn8c9513kulec_Y

Want to join the discussion live on Google Meet? Get a calendar invite by filling out this form: https://red.ht/office-hours

Timestamps:
00:00 vLLM intro
01:12 vLLM v0.23 project update
21:24 New vLLM course on DeepLearning.AI
26:03 vLLM Semantic Router deep dive, with live demo
54:03 Q&A
[vLLM Office Hours #52] - vLLM Semantic Router: Safer, Faster, Multi-Model Inference - June 25, 2026Secrets Management in Red Hat OpenShiftInfrastructure At The EdgeWhy should I standardize on RHEL?[vLLM Office Hours #55] - Mooncake + vLLM/llm-d Deep Dive - August 6, 2026GitOps Guide to the Galaxy (ep 108) | Declarative Networking w/IsovalentIn the Clouds (E52) | Elevating the Enterprise: Red Hat Summit 2026 Preview ft. Chuck DubuqueMonday highlights from Red Hat Summit 2026What’s New In Red Hat AI 3.4?The path from CentOS Linux to RHELAI Explained: Reduce GPU costs with LLM CompressorLearn about secure enclaves for AI inference
Red Hat |

[vLLM Office Hours #52] - vLLM Semantic Router: Safer, Faster, Multi-Model Inference - June 25, 2026

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER