SNIA SDCStorageAI 2026-Scaling Inference w/ KV Cache Storage Offload & RDMA Accelerated Architecture @SNIAVideo
SNIA SDCStorageAI 2026-Scaling Inference w/ KV Cache Storage Offload & RDMA Accelerated Architecture  @SNIAVideo
Uploaded May 2026 | Updated September 2026, 2 weeks ago
As LLMs become central to applications such as conversational AI, document processing, agentic workflows, and RAG, inference systems must support longer context windows and deeper interaction patterns. A major scalability limiter in these workloads is the KV Cache, which expands with context length and can rapidly exceed aggregated GPU memory capacity.
In this talk, we will focus on KV Cache and present how storage backed KV cache offloading leveraging high performance networked storage systems enables scalable inference beyond GPU memory limits. We highlight how RDMA enabled data paths and low latency storage significantly accelerate cached data movement, reducing end to end inference latency while supporting larger, more complex workloads.
We also share multi turn inference benchmarking results that expose the challenges of context accumulation in real world interaction in real world interaction sequences, including human computer multi turn interactions. As LLMs increasingly rely on iterative, multi step interactions, evaluating these real world workloads is essential for understanding system behavior and designing scalable inference architectures.

Presented by Ugur Kaynar | Dell - Distinguished Engineer, Storage Technologist – Storage CTO

Learn More:
• SDC: StorageAI Website: snia.org/sniadeveloper/storageai
• SNIA Website: snia.org
• SNIA Educational Library: snia.org/library
• X: twitter.com/SNIA
• LinkedIn: linkedin.com/company/snia
SNIA SDCStorageAI 2026-Scaling Inference w/ KV Cache Storage Offload & RDMA Accelerated ArchitectureScaling NGS Analyses for Ultra-High Complexity DNA Data Storage LibrariesHow New Memories Will Accelerate Both Inference and General Purpose ComputeWetlab validation of JPEG DNA coding system for image storage on  synthetic DNASNIA SDC 2025  - CXL Memory in WindowsSNIA SDC 2025  - Small Granularity Graph Neural Network Training & the Future of StorageBreaking the Memory Wall with MRAMThermodynamically favoured DNA data structures and algorithms
SNIAVideo |

SNIA SDCStorageAI 2026-Scaling Inference w/ KV Cache Storage Offload & RDMA Accelerated Architecture

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER