🤖 AI Summary
This study addresses the issue that reporting preferences acquired through Reinforcement Learning with Verifiable Rewards (RLVR) frequently conflict with users' immediate instructions, thereby degrading format adherence. Using the Qwen2.5-7B model and the GSM8K dataset, we conduct cross-evaluations and human intervention experiments to decouple model behavior into three independent dimensions: learned reporting preferences, instruction following, and answer correctness. Our findings reveal that although RLVR improves accuracy for specific formats, it substantially impairs generalization to alternative formats. Empirically, training with a boxed format causes response rates for hash-formatted outputs to plummet by 35%–74%, confirming that reporting preferences are learnable yet often contradict immediate instructions. These results highlight the inherent limitations of evaluation paradigms that rely exclusively on convention matching.
📝 Abstract
Reinforcement learning with verifiable rewards (RLVR) has become a prominent approach for improving language-model performance on reasoning tasks using automatically checked answers. Yet convention-matched evaluation cannot reveal whether reinforcing one reporting convention reduces adherence to a different request that the initial policy already follows. To test this, we train matched policies under two reporting conventions and evaluate each policy under both current requests, using the same initial policy as a shared reference. We complement this crossed design with controlled interventions and independent human calibration. On GSM8K, boxed-format RLVR reduces the fraction of Qwen2.5-7B responses containing the requested hash-format payload by 35.33--74.37 percentage points relative to a 95.45% initial baseline in four of five training seeds; the fifth improves by 2.50 points. In the four deteriorating runs, almost every response that omits the requested payload instead retains the trained boxed convention, and the same four seeds deteriorate under two fixed paraphrases. Changing only the final-answer marker in supervised targets reverses which reporting convention the model prefers across three seeds, providing controlled evidence that this preference is learnable. Across three settings with independent human calibration, gains under a convention-sensitive scorer exceed the corresponding gains in committed-answer correctness, i.e., the correctness of the answer the model actually commits to. Together, these results separate three distinct post-training outcomes: learned reporting preference, current-request adherence, and committed-answer correctness. They show that convention-matched accuracy alone does not fully characterize post-training behavior and motivate evaluating current-request adherence alongside convention-matched task accuracy.