The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the weakest-link failure in large model reasoning distillation caused by reward hacking and averaging constraints. We formulate distillation as reinforcement learning subject to worst-case constraints, proposing a non-augmented constrained MDP framework with low-variance policy gradient decomposition. Through GRPO optimization, KL regularization, and stateless augmented reward shaping, our method almost surely satisfies per-step log-likelihood lower bounds under penalty limits, circumventing the high computational cost of double Lagrangian optimization. Experiments demonstrate that this approach significantly expands the accuracy-fidelity Pareto frontier, matching pure RL accuracy while substantially reducing constraint violations, thereby achieving state-of-the-art strict reasoning success rates.
📝 Abstract
Distilling the reasoning capabilities of large language models (LLMs) into smaller students is a central challenge for efficient deployment. Current approaches face a fundamental tension: optimizing purely for verifiable task rewards (e.g., via GRPO) leads to reward hacking, where students arrive at correct final answers through flawed intermediate logic, while regularizing with soft divergence penalties against a teacher (e.g., KL-based distillation) dilutes task performance and, critically, allows the student to compensate for severe logical violations at one step with high teacher agreement at others. We argue that this averaging is fundamentally misaligned with the nature of reasoning: a chain-of-thought is only as valid as its weakest link. Motivated by this observation, we formulate reasoning distillation as a constrained reinforcement learning problem in which the task reward is maximized subject to a worst-case constraint on the teacher log-likelihood along every prefix of the trajectory. To avoid the prohibitive cost of dual Lagrangian solvers and the test-time teacher dependence of state-augmented methods such as Saute, we derive an unaugmented constrained MDP whose reward transformation preserves the hard-constraint semantics, admits a low-variance policy gradient decomposition into single-step and long-term terms, and provably satisfies the worst-case constraint almost surely in the penalty limit. Through extensive experiments on mathematical reasoning and code generation tasks, we demonstrate that our method significantly expands the accuracy-fidelity Pareto front. By matching the high Final Answer Correctness of pure RL and drastically reducing teacher constraint violations, we ultimately achieve the highest rigorous Reasoning Success Rate across all evaluated settings.
Problem

Research questions and friction points this paper is trying to address.

Reasoning Distillation
Reward Hacking
Worst-Case Constraint
Chain-of-Thought
Constrained Reinforcement Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Constrained Reinforcement Learning
Reasoning Distillation
Worst-Case Constraint
Reward Hacking
Unaugmented Constrained MDP
🔎 Similar Papers
No similar papers found.