NSDI 26 - RLBoost: Harvesting Preemptible Cloud Resources for Cost-Efficient Reinforcement Learning @UsenixOrg
NSDI 26 - RLBoost: Harvesting Preemptible Cloud Resources for Cost-Efficient Reinforcement Learning  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
RLBoost: Harvesting Preemptible Cloud Resources for Cost-Efficient Reinforcement Learning on LLMs

Yongji Wu, UC Berkeley; Xueshen Liu, University of Michigan; Haizhong Zheng, Carnegie Mellon University; Juncheng Gu, Google; Beidi Chen, Carnegie Mellon University; Z. Morley Mao, University of Michigan; Arvind Krishnamurthy, Google and University of Washington; Ion Stoica, UC Berkeley

Reinforcement learning (RL) has become essential for unlocking advanced reasoning capabilities in large language models (LLMs). RL workflows involve interleaving rollout and training stages with fundamentally different resource requirements. Rollout typically dominates overall execution time, yet scales efficiently through multiple independent instances. In contrast, training requires tightly-coupled GPUs with full-mesh communication. Existing RL frameworks fall into two categories: co-located and disaggregated architectures. Co-located frameworks fail to address this resource tension by forcing both stages to share the same GPUs. Disaggregated architectures, without modifications of well-established RL algorithms, suffer from resource under-utilization. Meanwhile, preemptible GPU resources, i.e., spot instances on public clouds and spare capacity in production clusters, present significant cost-saving opportunities for accelerating RL workflows, if efficiently harvested for rollout.

In this paper, we present RLBoost, a framework for cost-efficient RL training that harvests preemptible GPU resources. Our key insight is that rollout's stateless and embarrassingly parallel nature aligns perfectly with preemptible and often fragmented resources. To efficiently utilize these resources despite frequent and unpredictable availability changes, RLBoost adopts a hybrid architecture with three key techniques: (1) adaptive rollout offload to dynamically adjust workloads on the reserved (on-demand) cluster, (2) pull-based weight transfer that quickly provisions newly available instances, and (3) token-level response collection and migration for efficient preemption handling and continuous load balancing. Extensive experiments show RLBoost increases training throughput by 1.51x-1.97x while improving cost efficiency by 28%-49% compared to using only on-demand GPU resources. RLBoost is open-sourced at github.com/Terra-Flux/PolyRL.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - RLBoost: Harvesting Preemptible Cloud Resources for Cost-Efficient Reinforcement LearningPEPR 26 - CA-CI: A Normative Framework for Evaluating Privacy and Dignity in AI GovernanceNSDI 26 - Decoding RSSI Compression in RFID: Dynamic RCS Modeling and Tag-Intrinsic Power Metrics..NSDI 26 - Over-Threshold Multiparty Private Set Intersection for Collaborative...SREcon24 Europe/Middle East/Africa - Dude, You Forgot the Feedback: How Your Open Loop Control...NSDI 26 - Queue-Mem: Energy-Efficient Hardware Storage for Advanced Network Function AccelerationPEPR 26 - Dismantling the Barriers to Personal Data PortabilityNSDI 26 - Defending against Traffic Analysis Attacks with Flexible In-Network ObfuscationUSENIX Security 24 - HYPERPILL: Fuzzing for Hypervisor-bugs by Leveraging the Hardware...NSDI 26 - QCON: Seamless QoE-Aware 5G Streaming via Multi-ConnectivityPEPR 26 - Architecting Scalable Data Lineage Graph for Privacy Compliance and Agentic AnalysisFAST 26 - RosenBridge: A Framework for Enabling Express I/O Paths Across the Virtualization...
USENIX |

NSDI '26 - RLBoost: Harvesting Preemptible Cloud Resources for Cost-Efficient Reinforcement Learning

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER