Mitigating the Length-Scaling Tax with Online Distillation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the "length scaling tax" in reinforcement learning (RL) post-training, where verbose responses to already-solved queries degrade efficiency. To mitigate this issue, we propose length self-distillation, a method that requires no external teacher model. Instead, it constructs a teacher network via the exponential moving average of the policy to dynamically differentiate task difficulty: online distillation is applied to solved prompts to compress redundancy, while the RL objective is retained for unsolved prompts to preserve reasoning capabilities. Experimental results demonstrate that our approach matches or exceeds baseline performance while reducing the single-turn and multi-turn length scaling taxes to −3.7% and 13.7%, respectively. These findings indicate that the proposed method effectively curbs response inflation on simple queries without compromising overall model capability.
📝 Abstract
Length scaling during reinforcement-learning (RL) post-training is often viewed as a sign of improved reasoning ability, especially on difficult problems, but may also make responses to already-solved problems unnecessarily verbose. We quantify this side effect as the length-scaling tax (LST): excess response length on already-solved queries without a commensurate accuracy gain. To mitigate LST, we propose Length Self-Distillation (LSD), which routes solved prompts to on-policy distillation and retains the original RL objective for unsolved prompts. LSD uses an exponential moving average of the online policy as its teacher, requiring no external model. We find that LSD achieves comparable or better performance than RL across multiple variants, while substantially curbing response-length growth on easy queries. LSD reduces LST from 19.0% to -3.7% on single-turn reasoning and from 31.4% to 13.7% on multi-turn agentic tasks, demonstrating that LSD effectively preserves concise response patterns on easy queries while supporting efficient exploration on difficult queries during RL post-training.
Problem

Research questions and friction points this paper is trying to address.

Length-Scaling Tax
Reinforcement Learning
Post-training
Response Verbosity
Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Length-Scaling Tax
Length Self-Distillation
Online Distillation
Reinforcement Learning
Exponential Moving Average
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.