On Language Drift during RLVR Post-Training

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the causes of chain-of-thought (CoT) language drift following Reinforcement Learning with Verifiable Rewards (RLVR) post-training in large language models and its implications for monitorability. Through theoretical derivation and empirical analysis, we demonstrate for the first time that RLVR permits unbounded language drift whereas supervised fine-tuning (SFT) does not. We further establish that this phenomenon originates from training on novel tasks and cannot be constrained without sacrificing performance. Crucially, this work identifies an inherent trade-off between monitorability and reasoning capability, confirming that enhancing CoT monitorability in frontier models inevitably degrades their reasoning performance. These findings provide essential theoretical foundations for understanding RLVR mechanisms and optimizing safety alignment strategies.
📝 Abstract
Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)---have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed signs of language drift in their chains of thought (CoTs): unusual, non-standard, and seemingly nonsensical language use. Although it is well-documented---and can potentially impair CoT monitorability---the causes of language drift are thus far poorly understood. In this paper, we identify the conditions under which language drift occurs: we prove theoretically that RLVR optimization pressure permits unbounded language drift, while supervised fine-tuning does not. We then show empirically that language drift specifically arises during RLVR on novel reasoning tasks---i.e. when the target behavior cannot be drawn out of the base model. Finally, we prove that it is not possible to constrain language drift without constraining expected reward, suggesting that CoT monitorability cannot be improved without harming performance during RLVR post-training at the frontier.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Language Drift
RLVR
Chain-of-Thought
Reinforcement Learning
Monitorability
🔎 Similar Papers
No similar papers found.