NSDI 26 - HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds @UsenixOrg
NSDI 26 - HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds

Chiheng Lou, Sheng Qi, and Chao Jin, School of Computer Science, Peking University; Dapeng Nie, Haoran Yang, and Yu Ding, Alibaba Group; Xuanzhe Liu and Xin Jin, School of Computer Science, Peking University

With the proliferation of large language model (LLM) variants, developers are turning to serverless computing for cost-efficient LLM deployment. However, public cloud providers often struggle to provide performance guarantees for serverless LLM serving due to significant cold start latency caused by substantial model sizes and complex runtime dependencies. To address this problem, we present HydraServe, a serverless LLM serving system designed to minimize cold start latency in public clouds. HydraServe proactively distributes models across servers to quickly fetch them, and overlaps cold-start stages within workers to reduce startup latency. Additionally, HydraServe strategically places workers across GPUs to avoid network contention among cold-start instances. To minimize resource consumption during cold starts, HydraServe further introduces pipeline consolidation that can merge groups of workers into individual serving endpoints. Our comprehensive evaluations under diverse settings demonstrate that HydraServe reduces the cold start latency by 1.7×–4.7× and improves service level objective attainment by 1.43×–1.74× compared to baselines.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public CloudsNSDI 26 - Medley: Optimizing Midgress Bandwidth for Commercial Live Streaming CDNsNSDI 26 - Checkmate: Zero Performance Overhead Model Checkpointing via Network Gradient ReplicationPEPR 26 - Shadow Data in Tool Calls: The Privacy Leak Hiding in Plain SightSREcon26 Americas - The Zero Trust Odyssey: Our Journey to Modernize Internal AccessNSDI 26 - Geminet: Learning the Duality-based Topology-Agnostic Update Operator for Lightweight...NSDI 26 - Observability Is Eating Your Cores: Fine-Grained Analysis of Microservice Metrics with...NSDI 26 - UNUM: A New Framework for Network ControlPEPR 26 - Private AI: Building Trust Through Verifiable ComputationSREcon26 Americas - Building SRE Culture (without SREs, Technically)NSDI 26 - Detecting and Diagnosing Errors in Serving Archived Web PagesNSDI 26 - SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
USENIX |

NSDI '26 - HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER