VLLM K/V Caching With Ceph - Kyle Bader, IBM & Tushar Gohad, Intel @Cephstorage
VLLM K/V Caching With Ceph - Kyle Bader, IBM & Tushar Gohad, Intel  @Cephstorage
Uploaded November 2025 | Updated September 2026, 2 weeks ago
VLLM K/V Caching With Ceph - Kyle Bader, IBM & Tushar Gohad, Intel

Generative AI and LLMs are all the rage right now, and many people are asking where storage fits in and how it can help with either accelerating or reducing the cost of various AI workflows. In this session we will dive into a prototype Ceph caching plugin for vLLM that allows offloading attention states to Ceph, lowering the cost of inference by allowing clustered NVMe to complement GPU memory. We will describe how K/V caching fits into inference workloads, caching plugin implementation details, how we think it should evolve, and share some preliminary performance data.
VLLM K/V Caching With Ceph - Kyle Bader, IBM & Tushar Gohad, IntelBeyond Balancing: How Upmaps Were Used To Avoid Disaster! - Bryan Stillwell, Akamai TechnologiesSMB Keeping - Clients, Clusters & Clouds - John Mulligan, IBMIntroduction & State of the CephalopodCeph Developer Summit  - NVMe oF gatewayKeynote: Building Together: Open Storage for Today and Tomorrow - Joachim Kraftmayer, CEO, CLYSONVMeoF in Ceph - Whats New and Performance Best Practices - Mike Burkhart & Aviv Caro, IBMCeph CI and Developer Tooling: Past, Present, and FuturFastEC in Ceph Tentacle A Technical Deep Dive into Next Gen Erasure Coding PerformanceCephalocon 2025 Highlights – Vancouver’s Biggest Ceph Community Gathering!Ceph User + Dev Monthly Meetup - May 2026BXC - An Architecture for Lazy RBD Imaging and Snapshot Exports
Ceph |

VLLM K/V Caching With Ceph - Kyle Bader, IBM & Tushar Gohad, Intel

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER