NSDI 26 - FastServe: Iteration-Level Preemptive Scheduling for Large Language Model Inference @UsenixOrg
NSDI 26 - FastServe: Iteration-Level Preemptive Scheduling for Large Language Model Inference  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
FastServe: Iteration-Level Preemptive Scheduling for Large Language Model Inference

Bingyang Wu, Yinmin Zhong, Zili Zhang, Shengyu Liu, Fangyue Liu, Yuanhang Sun, Gang Huang, Xuanzhe Liu, and Xin Jin, School of Computer Science, Peking University

Large language models (LLMs) power a new generation of interactive AI applications exemplified by ChatGPT. The interactive nature of these applications demands low latency for LLM inference. Existing LLM serving systems use run-tocompletion processing for inference jobs, which suffers from head-of-line blocking and long latency. We present FastServe, a distributed LLM serving system which exploits the autoregressive pattern of LLM inference to enable preemption at the granularity of each output token. FastServe uses preemptive scheduling to minimize latency with a novel skip-join Multi-Level Feedback Queue scheduler. Based on the new semi information-agnostic setting of LLM inference, the scheduler leverages the input length information to assign an appropriate initial queue for each arrival job to join. Queues with higher priority than the one the job joins are skipped to reduce demotions. We design an efficient GPU memory management mechanism that proactively offloads and uploads intermediate state between GPU memory and host memory for LLM inference. Evaluation shows that compared to the state-of-the-art solution vLLM, FastServe improves the throughput by up to 6.1×.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - FastServe: Iteration-Level Preemptive Scheduling for Large Language Model InferenceNSDI 26 - Ubers Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale...SREcon26 Americas - Taming the Unpredictable: Reliability in ChaosNSDI 26 - RollPacker: Taming Long-Tail Rollouts for RL Post-Training with Tail BatchingNSDI 26 - KeepON: Supporting Deterministic Traffic on Standard NICsNSDI 26 - FENIX: Enabling In-Network DNN Inference with FPGA-Enhanced Programmable SwitchesComputer Security and Voting, Invited Talk by David Dill at USENIX Security 07PEPR 26 - Toward Provably Private Insights into AI UsePEPR 26 - DPSynth: From Research to Production—Engineering Differentially Private Synthetic...NSDI 26 - FRCC: Towards Provably Fair and Robust Congestion ControlUSENIX Security 25 - ALERT: Machine Learning-Enhanced Risk Estimation for Databases Supporting...NSDI 26 - cc-pipe: Breaking Systemic Bottlenecks in RPKI Data Supply Chain with Concurrent and...
USENIX |

NSDI '26 - FastServe: Iteration-Level Preemptive Scheduling for Large Language Model Inference

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER