PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the high multi-turn reasoning costs, low GPU utilization, and rigid configurations inherent in agent-based reinforcement learning by proposing a dynamic optimization framework that couples elastic resources with prefill-decode (PD) configuration. This work pioneers a resource elasticity and PD coordination mechanism leveraging unified state management, runtime performance prediction, cost-aware switching, and incremental migration strategies to accelerate asynchronous inference using idle training GPUs while minimizing reconfiguration overhead. Experimental results demonstrate that the proposed framework improves system throughput by 2.17Γ— to 2.79Γ— compared to fixed-resource systems and achieves up to a 36.3% speedup over RLBoost+.
πŸ“ Abstract
Multi-turn rollout dominates the cost of agentic reinforcement learning (RL). Asynchronous execution and elastic GPU resources can accelerate this stage, but adding rollout replicas yields diminishing returns while training GPUs remain idle between updates. We observe that effective resource use also depends on the prefill--decode (PD) configuration. Both the choice between colocation and disaggregation and the optimal PD ratio vary with the workload, making resource scaling and PD configuration interdependent. Exploiting this opportunity requires selecting effective configurations and realizing their benefits within transient resource-availability windows despite reconfiguration costs. We present PEARL, an asynchronous agentic RL system that coordinates external resource elasticity, temporary reuse of idle training GPUs, and adaptive PD execution. PEARL maintains a unified GPU--worker--role state and uses runtime profiles to predict rollout batch completion time, accounting for environment-induced reductions in decode concurrency. It selects the PD mode and ratio under the current GPU budget and translates each decision into an incremental transition plan that minimizes worker and role changes. Cost-aware switching and borrowing policies suppress transitions with insufficient expected benefit while ensuring timely return of training GPUs. Our evaluation show that PEARL achieves $2.17$--$2.79\times$ the throughput of fixed-resource ROLL across different LLMs. Compared with RLBoost+, throughput improves by up to approximately 26.9\% for Qwen3-8B and 36.3\% for Qwen3-30B-A3B.
Problem

Research questions and friction points this paper is trying to address.

Agentic Reinforcement Learning
Multi-turn Rollout
Prefill-Decode Configuration
Elastic Resource Scheduling
GPU Utilization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Reinforcement Learning
Prefill-Decode Disaggregation
Elastic Resource Scheduling
Asynchronous Execution
Cost-aware Switching
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Jiaan Zhu
Jiaan Zhu
δΈ­ε›½η§‘ε­¦ζŠ€ζœ―ε€§ε­¦
Computer ScienceLarge Language Model
W
Wei Gao
HKUST
Y
Youhui Bai
USTC
Zewen Jin
Zewen Jin
University of Science and Technology of China
LLM Training / ServingMoEServerless Computing
J
Ju Huang
Alibaba Group
S
Siran Yang
Alibaba Group
J
Jiamang Wang
Alibaba Group
L
Lin Qu
Alibaba Group
C
Cheng Li
USTC