Quality Determines Direction, Length Shapes Magnitude: Length Control for Open-Ended Reinforcement Learning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency in open-ended reinforcement learning caused by the entanglement of response length and quality, proposing the Quality-Gated Length Advantage Shaping (QGLAS) algorithm. QGLAS follows an asymmetric principle—“quality determines direction, length modulates magnitude”—by first computing advantages based on quality, then applying bounded rewards exclusively to short, positive samples, and introducing adaptive intra-group quality separation adjustment. The core innovation lies in decoupling the effects of length and quality through a quality-gating mechanism, effectively preventing reward shaping from distorting advantage signs or magnitudes. Experiments demonstrate that under a 30% compression rate, QGLAS retains 98.4%–102.0% of quality gains, significantly outperforming the 68.3%–75.5% achieved by baseline methods, thereby enabling efficient length control.
📝 Abstract
Reinforcement learning (RL) changes not only what language models say, but also how much they say, often increasing response length at the cost of token efficiency. Controlling this length growth is particularly challenging in open-ended RL because (i) response length is entangled with quality, (ii) open-ended tasks lack a natural success boundary for deciding when efficiency should be prioritized, and (iii) dense, graded rewards often yield small within-group quality margins, making quality-induced advantages especially sensitive to reward-level length shaping, which can perturb their magnitudes and even reverse their signs. We therefore adopt an asymmetric principle: quality should determine the direction of reinforcement, while length should only shape its magnitude. We instantiate this principle with Quality-Gated Length Advantage Shaping (QGLAS), which first computes advantages from quality rewards alone, then adds bounded bonuses only to shorter positive-advantage responses, leaving all other advantages unchanged. The bonus strength is further adapted to within-group quality separation, allowing conciseness to matter more when quality-favored responses are similar and less when their quality differences are clear. Across different model families, open-ended benchmarks, and reward sources, QGLAS consistently achieves a stronger quality--length trade-off than representative baselines. At approximately 30% compression, QGLAS retains 98.4--102.0% of the macro-average quality gains achieved by quality-only RL over the base model, compared with 68.3--75.5% for these baselines at comparable compression.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Open-Ended Reinforcement Learning
Length Control
Quality-Gated Length Advantage Shaping (QGLAS)
Token Efficiency
Advantage Shaping
🔎 Similar Papers
No similar papers found.
Z
Zijun Weng
1Fudan University; 2Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd.
Z
Zhongan Bi
3Zhejiang University
X
Xuanang Gao
4Shanghai Jiao Tong University
X
Xiaohui Hu
2Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd.
S
Shuangyong Song
2Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd.
Yongxiang Li
Yongxiang Li
Professor, RMIT University
Electronic Materials and Devices
K
Kaidong Yu
2Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd.
X
Xuanjing Huang
1Fudan University