SNIA SDC 2025  - Disaggregated KV Storage: A New Tier for Efficient Scalable LLM Inference @SNIAVideo
SNIA SDC 2025  - Disaggregated KV Storage: A New Tier for Efficient Scalable LLM Inference  @SNIAVideo
Uploaded November 2025 | Updated September 2026, 2 weeks ago
As generative AI models continue to grow in size and complexity, the infrastructure costs of inference—particularly GPU memory and power consumption—have become a limiting factor. This session presents a disaggregated key-value (KV) storage architecture designed to offload KV-cache tensors efficiently, reducing GPU compute pressure while maintaining low-latency, high-throughput inference. We introduce the first end-to-end system based on shared storage for KV-cache offloading. It integrates with production-scale orchestration frameworks such as Dynamo and Production Stack, enabling scalable deployment across distributed GPU clusters. We provide both theoretical analysis and empirical evaluation, comparing our approach to state-of-the-art inference engines such as vLLM. Our benchmarks demonstrate 5–8× higher request throughput and 5–7× faster prefill latency compared to baseline systems. Experiments cover a range of GPU types and LLMs, including DeepSeek-V3, and simulate diverse use cases such as multi-turn conversations, long context generation, and agentic workloads. Unlike traditional block or file storage systems, which are not optimized for the fine-grained, high-frequency access patterns of LLM workloads, our stateless external KV store enables direct GPU-initiated I/O and overlapping of compute and data access, improving efficiency at the infrastructure level. This session will provide technical insights into system design, performance characteristics, and practical deployment lessons. It is intended for engineers, system architects, and infrastructure practitioners seeking scalable, storage-centric approaches to improve the efficiency and elasticity of LLM inference at scale.
Learn how disaggregated KV storage enhances GPU utilization and reduces compute cycles, specifically, reduce time to generate first token and time between tokens in popular inference engines such as vLLM Explore how stateless KV stores support elastic scaling across GPU clusters: specifically, simplify load balancing, offloads index and placement management Gain practical insights into infrastructure-level details to support optimizations of SOTA inference frameworks such as DeepSeek and Dynamo like MLA, prefill-decode disaggregation Gain strategies to reduce CPU-GPU synchronization overhead data movement optimizations that improve throughput and latency in AI pipelines, specifically, GPU-initiated and GPU-triggered IO.
Presented by Eshcar Hillel, Pliops
Learn More:
• SDC Website: snia.org/sniadeveloper
• SNIA Website: snia.org
• SNIA Educational Library: snia.org/library
• X: twitter.com/SNIA
• LinkedIn: linkedin.com/company/snia
SNIA SDC 2025  - Disaggregated KV Storage: A New Tier for Efficient Scalable LLM InferenceSNIA SDC 2025  - CSAL w/ Core Scaling for RAID5F: Revolutionizing Cloud Storage Perf & ReliabilityShort Video: SAS Continues to Innovate with SBC-5SNIA Partner Update - CXL - June 20242026 SNIA Preview Welcome and IntroductionSNIA SDC 2025  - Global Distributed Client-side CacheSNIA Technical Council - 2026 SNIA PreviewSNIA SDC 2025  - Highly Scalable, Masterless, Distributed File System at RubrikSNIA SDC 2025  - Storage for AI 102SNIA SDC 2025  - SMR and HAMR Advancing HDD Areal DensityThe Next Phase of DNA Data Storage: Building an Industry EcosystemSNIA SDC 2025  -MEXT Radical Reduction in Computing Costs
SNIAVideo |

SNIA SDC 2025 - Disaggregated KV Storage: A New Tier for Efficient Scalable LLM Inference

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER