NSDI 26 - SYMPHONY: Enabling Compute-Memory Disaggregation in LLM Serving Systems @UsenixOrg
NSDI 26 - SYMPHONY: Enabling Compute-Memory Disaggregation in LLM Serving Systems  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
NSDI '26 - SYMPHONY: Enabling Compute-Memory Disaggregation in LLM Serving Systems

Saurabh Agarwal and Bodun Hu, UT-Austin; Anyong Mao, UW-Madison; Aditya Akella, UT-Austin; Shivaram Venkataraman, UW-Madison

Large Language Models (LLMs) power AI applications such as chatbots and agents, which maintain conversational state across multiple turns. Serving these workloads is inherently stateful: each request generates a KV cache storing token-level state. Existing systems either recompute caches or offload them to host memory—both approaches incur high latency, cause load imbalance, and limit scalability. We present SYMPHONY, a disaggregated memory management layer that decouples compute from KV cache storage while meeting strict latency requirements. To enable disaggregation, SYMPHONY employs advisory requests—prefetching hints derived from user interactions or workload structure—to move caches off the critical path and enable fine-grained, request-level load balancing. Since these predictive signals are often unreliable, SYMPHONY introduces two key techniques: priority-based KV cache management, which allocates memory based on neural network structure and request priority, and cooperative memory management, which dynamically coordinates GPU memory with the serving framework. Evaluations on LLaMA models with ShareGPT and Burst-GPT workloads show that SYMPHONY reduces end-to-end latency by 2.4× over vLLM and serves 4× more requests with minimal latency increase.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - SYMPHONY: Enabling Compute-Memory Disaggregation in LLM Serving SystemsPEPR 26 - Provenance Without Surveillance: Privacy Engineering for AI Content TransparencyNSDI 26 - A Fast Solver-Free Algorithm for Traffic Engineering in Large-Scale Data Center NetworkNSDI 26 - Building A CSFQ-Inspired Transport for Switched CXL Memory PoolingSREcon24 Europe/Middle East/Africa - Noisy Neighbors, through NetworkingNSDI 26 - Sparse Checkpointing for Fast and Reliable MoE TrainingVehicleSec 25 - CarPlay at Risk: Unveiling Security Threats of Third-Party Infotainment AdaptersPEPR 26 - The Emperors New Embeddings: Obfuscating ML Inputs Doesnt Provide PrivacyNSDI 26 - RLBoost: Harvesting Preemptible Cloud Resources for Cost-Efficient Reinforcement LearningPEPR 26 - CA-CI: A Normative Framework for Evaluating Privacy and Dignity in AI GovernanceNSDI 26 - Decoding RSSI Compression in RFID: Dynamic RCS Modeling and Tag-Intrinsic Power Metrics..NSDI 26 - Over-Threshold Multiparty Private Set Intersection for Collaborative...
USENIX |

NSDI '26 - SYMPHONY: Enabling Compute-Memory Disaggregation in LLM Serving Systems

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER