NSDI 26 - RollPacker: Taming Long-Tail Rollouts for RL Post-Training with Tail Batching @UsenixOrg
NSDI 26 - RollPacker: Taming Long-Tail Rollouts for RL Post-Training with Tail Batching  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
RollPacker: Taming Long-Tail Rollouts for RL Post-Training with Tail Batching

Wei Gao, Yuheng Zhao, Dakai An, Tianyuan Wu, and Lunxi Cao, Hong Kong University of Science and Technology; Shaopan Xiong, Ju Huang, Weixun Wang, Siran Yang, Wenbo Su, Jiamang Wang, Lin Qu, and Bo Zheng, Alibaba Group; Wei Wang, Hong Kong University of Science and Technology

Reinforcement Learning (RL) is a pivotal post-training technique for enhancing the reasoning capabilities of Large Language Models (LLMs). However, synchronous RL post‑training frequently suffers from significant GPU underutilization—often referred to as pipeline "bubbles"—caused by imbalanced response lengths within rollout steps. Many RL systems attempt to alleviate this problem by relaxing synchronization, but this can compromise training accuracy.

In this paper, we introduce tail batching, a novel rollout scheduling strategy for synchronous RL. Tail batching systematically consolidates prompts leading to long-tail responses into a few designated "long rounds", ensuring that the majority of rollout steps ("short rounds") contain only balanced, short responses. By strategically reordering execution, this approach dramatically reduces GPU idle time and accelerates RL training without sacrificing on-policy accuracy. We present RollPacker, a system that fully harnesses the benefits of tail batching through holistic optimizations across all three RL stages: elastic parallelism adaptation for rollout, dynamic resource allocation and scheduling for reward, and stream-based training. Cluster deployment on up to 128 H800 GPUs demonstrates that RollPacker achieves an end-to-end training speedup of 2.03× to 2.56× over veRL, and up to 2.24× speedup compared to RLHFuse across the Qwen2.5 family of LLMs. The code is available at github.com/Farrrrland/RollPacker.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - RollPacker: Taming Long-Tail Rollouts for RL Post-Training with Tail BatchingNSDI 26 - KeepON: Supporting Deterministic Traffic on Standard NICsNSDI 26 - FENIX: Enabling In-Network DNN Inference with FPGA-Enhanced Programmable SwitchesComputer Security and Voting, Invited Talk by David Dill at USENIX Security 07PEPR 26 - Toward Provably Private Insights into AI UsePEPR 26 - DPSynth: From Research to Production—Engineering Differentially Private Synthetic...NSDI 26 - FRCC: Towards Provably Fair and Robust Congestion ControlUSENIX Security 25 - ALERT: Machine Learning-Enhanced Risk Estimation for Databases Supporting...NSDI 26 - cc-pipe: Breaking Systemic Bottlenecks in RPKI Data Supply Chain with Concurrent and...NSDI 26 - Who Watches the Watchers? On the Reliability of Softwarizing Cloud Application ManagementNSDI 26 - SLATE: Service Layer Traffic EngineeringPEPR 26 - Scaling Privacy Threat Modeling: From Architects to Developers
USENIX |

NSDI '26 - RollPacker: Taming Long-Tail Rollouts for RL Post-Training with Tail Batching

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER