NSDI 26 - DroidSpeak: KV Cache Sharing Across Fine-tuned Model Variants @UsenixOrg
NSDI 26 - DroidSpeak: KV Cache Sharing Across Fine-tuned Model Variants  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
DroidSpeak: KV Cache Sharing Across Fine-tuned Model Variants

Yuhan Liu, Yuyang Huang, Jiayi Yao, Shaoting Feng, Zhuohan Gu, Kuntai Du, Hanchen Li, Yihua Cheng, and Junchen Jiang, University of Chicago; Shan Lu, Madan Musuvathi, and Esha Choukse, Microsoft

Compound AI systems, such as agentic systems, are an emerging trend in large-scale enterprise settings, with multiple LLMs specialized for different users, tasks, and/or roles working together. In these scenarios, different models often process inputs that share the same context prefix. Although much work was done in the past to enable the reuse of prefix KV caches across inputs for a single model, how to enable one model to reuse the prefix KV caches of a different model remains an open question.

We introduce DroidSpeak, the first distributed LLM inference system that enables KV cache reuse across distributed nodes running inference of different LLMs, so long as the LLMs have the same architecture. We present the first study that aims at understanding the impact of sharing KV caches across different LLMs, and if/when such sharing affects quality. Inspired by the findings, we present DroidSpeak, which selectively recomputes a few layers of the KV cache produced by another LLM and reuses the remaining layers, with negligible quality loss. Moreover, carefully pipelining the layer-wise re-computation and the loading of reused KV cache further improves the inference performance. Experiments on diverse datasets and model pairs demonstrate that DroidSpeak achieves up to 4x throughput improvement and about 3.1× faster prefill (time to first token), with negligible loss of quality in F1 scores, Rouge-L or code similarity score, compared to the baseline which does not allow any sharing across models.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - DroidSpeak: KV Cache Sharing Across Fine-tuned Model VariantsNSDI 26 - REAL: Emulating Control Plane at Simulator’s CostPEPR 26 - Enforcement of Data Protection Laws in Africa: Implications for Privacy EngineersPEPR 26 - Adopting AI in Local Government with Privacy and Equity in Mind: A Case Study of the...SREcon26 Americas - Stop Reading Changelogs: Safer Kubernetes Upgrades with SimulationNSDI 26 - Skyline: A Cloud Centric Internet Monitoring EngineNSDI 26 - HybridMesh: A Hardware-software Hybrid Approach for Accelerating Service Mesh IngressSREcon26 Americas - AI Agents for Incident Investigation: The Good, The Bad, and The UglySREcon26 Americas - 5 Wrong Hypotheses about PostgreSQL Multi-Transaction LocksSREcon25 Europe/Middle East/Africa - Utilization Is the Key to Efficiency: What It Takes to Run...NSDI 26 - PolicyCache: Intra-flow Learning in Congestion ControlNSDI 26 - From Source to Solution: Tackling Packet Losses in Large-scale Cloud Gaming...
USENIX |

NSDI '26 - DroidSpeak: KV Cache Sharing Across Fine-tuned Model Variants

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER