Where Does Staleness Accumulate? Pool Aware Effective Staleness Control for Asynchronous RL in LLM Post-Training

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation caused by policy lag in asynchronous reinforcement learning for large language models (LLMs) by proposing PACE. This method pioneers the decoupling of trajectory staleness into generation and waiting components, employing a prefix-aware effective staleness score to prevent the erroneous penalization of long trajectories. By integrating an adaptive rejection budget to optimize training pool management, PACE achieves precise staleness control without wall-clock overhead. Experimental results demonstrate that PACE improves single-turn inference accuracy by 18.7% while reducing GPU time by 47.1%. Furthermore, it significantly outperforms both synchronous and unfiltered asynchronous baselines in multi-turn tool reasoning and mixture-of-experts (MoE) models, effectively balancing high resource utilization with training stability.
📝 Abstract
Fully asynchronous reinforcement learning (RL) improves resource utilization in large language model post-training by overlapping rollout generation with policy optimization, but it also introduces policy lag as trajectories are generated and queued while the trainer continues to update. We study how this lag accumulates over a trajectory's lifetime and how it can be controlled without sacrificing the wall-clock benefits of asynchronous execution. We decompose trajectory staleness into Generation Staleness, accumulated before rollout completion, and Waiting Staleness, accumulated after a completed trajectory enters the pool. Motivated by this decomposition, we introduce PACE (Pool-Aware Control of Effective Staleness). PACE converts excess pool occupancy into an adaptive rejection budget and ranks completed trajectories using an effective-staleness score that combines Waiting Staleness with prefix-aware Generation Staleness. This avoids penalizing long or interrupted rollouts solely because they span multiple policy versions. In single-turn mathematical reasoning, PACE improves the six-benchmark average validation accuracy by 18.7\% over unfiltered asynchronous RL at the same wall-clock budget and matches synchronous RL performance with 47.1\% less GPU time. PACE also improves validation performance in multi-turn tool-integrated reasoning, outperforming both synchronous and unfiltered asynchronous RL. Further experiments with the mixture-of-experts model and an alternative RL algorithm support its applicability across model architectures and training algorithms.
Problem

Research questions and friction points this paper is trying to address.

Asynchronous Reinforcement Learning
LLM Post-Training
Staleness
Policy Lag
Trajectory Pool
Innovation

Methods, ideas, or system contributions that make the work stand out.

Asynchronous Reinforcement Learning
Staleness Control
Trajectory Pool
LLM Post-Training
Effective Staleness Score