NSDI 26 - FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees @UsenixOrg
NSDI 26 - FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees

Gabriele Oliaro, Carnegie Mellon University; Xupeng Miao, Purdue University; Xinhao Cheng, Carnegie Mellon University; Vineeth Kada, Anthropic PBC; Mengdi Wu, Ruohan Gao, and Yingyi Huang, Carnegie Mellon University; Remi Delacourt, Mistral AI; April Yang, Carnegie Mellon University; Yingcheng Wang, Purdue University; Colin Unger, Stanford University; Zhihao Jia, Carnegie Mellon University and Amazon Web Services

Finetuning large language models (LLMs) is essential for task adaptation, yet today's serving stacks isolate inference and finetuning on separate GPU clusters—wasting resources and under-utilizing hardware. We introduce FlexLLM, the first system to co-serve LLM inference and PEFT-based finetuning on shared GPUs by fusing computation at the token level. FlexLLM's static compilation optimizations—dependent parallelization and graph pruning significantly shrink activation memory, leading to end-to-end GPU memory savings by up to 80%. At runtime, a novel token-level finetuning mechanism paired with a hybrid token scheduler dynamically interleaves inference and training tokens within each co-serving iteration, meeting strict latency SLOs while maximizing utilization. In end-to-end benchmarks on LLaMA-3.1-8B, Qwen-2.5-14B, and Qwen-2.5-32B, FlexLLM maintains inference SLO compliance at up to 20 req/s, and improves finetuning throughput by 1.9-4.8× under heavy inference workloads and 2.5-6.8× under light loads, preserving over 76% of peak finetuning progress even at peak demand. FlexLLM is publicly available at https://flexllm.github.io.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO GuaranteesNSDI 26 - Libra: Flexible Request Partitioning and Scheduling for Serving Unbalanced and Dynamic...SREcon26 Americas - The Ironies of AI²USENIX Security 25 - Serverless Functions Made Confidential and Efficient with Split ContainersNSDI 26 - Enabling AI Network Cross-Layer Design and Operations with Arcadia...NSDI 26 - KRAKENGUARD: Towards Fine-Grained eBPF IsolationNSDI 26 - Eywa: Automating Model-Based Testing using LLMsSREcon26 Americas - Executing Chaos Engineering in Production at a Critical Financial InstitutionSREcon26 Americas - The Case of the Misnamed Cities: CAST Analysis of a Google Maps IncidentNSDI 26 - Attack of the Bubbles: Straggler-Resilient Pipeline Parallelism for Large Model TrainingNSDI 26 - Matryoshka: Realizing Hyperscale Data Center Network Design for the AI EraSREcon26 Americas - Low Latency Serving of Offline Data: Efficient, Safe, and Reliable Data...
USENIX |

NSDI '26 - FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER