NSDI 26 - Libra: Flexible Request Partitioning and Scheduling for Serving Unbalanced and Dynamic... @UsenixOrg
NSDI 26 - Libra: Flexible Request Partitioning and Scheduling for Serving Unbalanced and Dynamic...  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
Libra: Flexible Request Partitioning and Scheduling for Serving Unbalanced and Dynamic LLM Workloads

Chaoyi Ruan, National University of Singapore; Yinhe Chen, Dongqi Tian, and Yandong Shi, University of Science and Technology of China; Yongji Wu, UC Berkeley; Jialin Li, National University of Singapore; Cheng Li, University of Science and Technology of China and Institute of Artificial Intelligence, Hefei Comprehensive National Science Center

LLM inference must meet strict latency SLOs while maximizing throughput. Yet, real-world variability in prompt and response lengths skews compute-intensive prefill and memory-bound decode phases, making both colocated (even with chunked prefill) and disaggregated deployments unable to simultaneously deliver low tail latency and high throughput.

We introduce Libra, a high performance LLM serving system that maximizes goodput under SLO constraints even when handling imbalanced and dynamic workloads. At the core of Libra is a micro-request based flexible partitioning and scheduling (FPS) abstraction. The abstraction splits each request at any token boundary into multiple cooperating segments. Libra then designs a two-level scheduling framework that balances micro-request load across unified GPU instances. The framework consists of a global scheduler that selects per-request split points, and a local scheduler on each GPU instance to form SLO-aware batches. Finally, Libra uses chunked KV cache transfers to support cross-instance micro-request execution. On real-world traces, Libra improves goodput by up to 1.91× and 1.61×, increases serving capacity from 1.15× to 3.07×, and improves serving performance by up to 74.2% in a hybrid workload under strict SLOs and A100/H100 GPUs compared to state-of-the-art colocated and disaggregated baselines.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - Libra: Flexible Request Partitioning and Scheduling for Serving Unbalanced and Dynamic...SREcon26 Americas - The Ironies of AI²USENIX Security 25 - Serverless Functions Made Confidential and Efficient with Split ContainersNSDI 26 - Enabling AI Network Cross-Layer Design and Operations with Arcadia...NSDI 26 - KRAKENGUARD: Towards Fine-Grained eBPF IsolationNSDI 26 - Eywa: Automating Model-Based Testing using LLMsSREcon26 Americas - Executing Chaos Engineering in Production at a Critical Financial InstitutionSREcon26 Americas - The Case of the Misnamed Cities: CAST Analysis of a Google Maps IncidentNSDI 26 - Attack of the Bubbles: Straggler-Resilient Pipeline Parallelism for Large Model TrainingNSDI 26 - Matryoshka: Realizing Hyperscale Data Center Network Design for the AI EraSREcon26 Americas - Low Latency Serving of Offline Data: Efficient, Safe, and Reliable Data...PEPR 26 - The Disposable Identity: Eliminating Non-Human Identity Risk in Federal Healthcare...
USENIX |

NSDI '26 - Libra: Flexible Request Partitioning and Scheduling for Serving Unbalanced and Dynamic...

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER