QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the training inefficiencies caused by GPU idleness and trajectory redundancy in online reinforcement learning for ultra-long-horizon LLM agents. To this end, we propose an end-to-end optimization framework that incorporates a non-disruptive elastic GPU scheduling mechanism to enable dynamic resource reallocation. Furthermore, the framework performs trajectory deduplication by integrating partial progress scoring with a nonlinear branching history reconstruction algorithm. Experimental results demonstrate that the proposed approach achieves an absolute performance gain of 6.0% on the NL2RepoBench benchmark and accelerates training by up to 1.85× compared to state-of-the-art methods.
📝 Abstract
Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental challenges: (1) severe execution variance and prolonged rollout delays cause massive GPU idling; and (2) complex non-linear branching generates massive trajectory redundancy, crippling training efficiency. To address these, we presents QwenGyre, an end-to-end framework for xlong-horizon online RL. QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs. Scaled to our flagship model, Qwen~3.8 2.4T, with 700K tokens per rollout, QwenGyre yields a 6.0% absolute gain on NL2RepoBench (52.5% $\to$ 58.5%) in 48 steps. Across our evaluations on diverse domains of training datasets, QwenGyre delivers up to $1.85\times$ and $1.78\times$ speedups over Colocate and Async, respectively.
Problem

Research questions and friction points this paper is trying to address.

xLong-Horizon Agents
Online Reinforcement Learning
GPU Idling
Trajectory Redundancy
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Elastic Reinforcement Learning
xLong-Horizon Agents
Trajectory Deduplication
Online RL
GPU Reallocation