🤖 AI Summary
This study addresses the misalignment between cross-entropy loss and verifier objectives during the post-training of large language models on verifiable tasks, which causes probability mass to diffuse toward incorrect outputs and traps models in suboptimal policies. By identifying this failure mode within imitation learning, the authors introduce Shannon entropy as a differentiable proxy for support set size and propose an entropy-regularized cross-entropy method. This approach integrates a token-level Shannon entropy regularizer with the standard cross-entropy loss to effectively constrain the policy support set, preventing probability shifts toward erroneous outputs. Evaluated on mathematical reasoning and code generation benchmarks, the proposed method significantly improves verifier accuracy, establishing a superior optimization paradigm for the post-training of large language models.
📝 Abstract
Large language models are often post-trained on expert demonstrations using cross-entropy (CE), even when the downstream objective is not to imitate the demonstrated solution but to produce any output accepted by a verifier. This mismatch is seen in verifiable domains with multiple correct solutions, such as mathematical reasoning and code generation, where training data may contain only one expert solution per problem. We show that minimizing cross-entropy can be misaligned with minimizing verifier risk; two policies can assign identical likelihood to the observed demonstrations while placing different probability mass on incorrect outputs. This is formalized through a learning-theoretic counterexample in which CE minimization selects a suboptimal policy. We identify that controlling the support of the learned policy can solve this problem by preventing probability mass from spreading to unsupported outputs. Since support size is non-differentiable and computationally intractable, we propose entropy-regularized cross-entropy (ER-CE), using token-level Shannon entropy as a tractable proxy. Finally, across mathematical reasoning and code-generation benchmarks, we find that entropy-regularized training consistently improves verifier accuracy over standard cross-entropy. Our results identify a simple failure mode of imitation-based post-training in verifiable tasks and provide a practical objective that is better aligned with producing correct outputs.