Entropy Regularization: A Free Correction to Cross-Entropy for Verified Demonstrations

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misalignment between cross-entropy loss and verifier objectives during the post-training of large language models on verifiable tasks, which causes probability mass to diffuse toward incorrect outputs and traps models in suboptimal policies. By identifying this failure mode within imitation learning, the authors introduce Shannon entropy as a differentiable proxy for support set size and propose an entropy-regularized cross-entropy method. This approach integrates a token-level Shannon entropy regularizer with the standard cross-entropy loss to effectively constrain the policy support set, preventing probability shifts toward erroneous outputs. Evaluated on mathematical reasoning and code generation benchmarks, the proposed method significantly improves verifier accuracy, establishing a superior optimization paradigm for the post-training of large language models.
📝 Abstract
Large language models are often post-trained on expert demonstrations using cross-entropy (CE), even when the downstream objective is not to imitate the demonstrated solution but to produce any output accepted by a verifier. This mismatch is seen in verifiable domains with multiple correct solutions, such as mathematical reasoning and code generation, where training data may contain only one expert solution per problem. We show that minimizing cross-entropy can be misaligned with minimizing verifier risk; two policies can assign identical likelihood to the observed demonstrations while placing different probability mass on incorrect outputs. This is formalized through a learning-theoretic counterexample in which CE minimization selects a suboptimal policy. We identify that controlling the support of the learned policy can solve this problem by preventing probability mass from spreading to unsupported outputs. Since support size is non-differentiable and computationally intractable, we propose entropy-regularized cross-entropy (ER-CE), using token-level Shannon entropy as a tractable proxy. Finally, across mathematical reasoning and code-generation benchmarks, we find that entropy-regularized training consistently improves verifier accuracy over standard cross-entropy. Our results identify a simple failure mode of imitation-based post-training in verifiable tasks and provide a practical objective that is better aligned with producing correct outputs.
Problem

Research questions and friction points this paper is trying to address.

cross-entropy
verifier risk
large language models
post-training
misalignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Entropy Regularization
Cross-Entropy
Verifiable Tasks
Support Control
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Mihir Dhanakshirur
Mihir Dhanakshirur
Indian Institute of Science
Statistical LearningMachine LearningCausal Inference
A
Adam Ousherovitch
Department of Statistics, University of Michigan
A
Ambuj Tewari
Department of Statistics, University of Michigan