ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue that structured pruning degrades the reasoning capabilities of models, thereby limiting the effectiveness of recovery through online policy distillation. To overcome this, we propose Recovery-Aware Calibration (ReCal), a method that pioneers shifting recovery considerations to the pruning stage. Specifically, ReCal computes the forward KL divergence between an unpruned teacher and pruned probes to identify perturbed predictions, subsequently reweighting calibration statistics to refine existing pruning criteria and guide the model in preserving critical predictions. This plug-and-play approach significantly mitigates residual damage from pruning. It yields performance gains of up to 16.7 percentage points on the AIME mathematical reasoning benchmark while concurrently improving code generation, effectively enhancing the distillation-based recovery capacity of pruned models.
📝 Abstract
Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery. Because OPD relies on student-generated trajectories, pruning damage that persists after offline distillation can limit its effectiveness. We propose RECAL, Recovery-Aware Calibration, a simple plug-and-play approach that improves OPD recovery by adjusting calibration before pruning. RECAL uses forward KL between an unpruned teacher and a pruned probe to identify teacher-supported predictions disrupted by pruning, then reweights calibration statistics to guide existing pruning criteria toward preserving these predictions. Across multiple models and pruning methods, RECAL consistently improves mathematical reasoning after OPD, achieving gains of up to 16.7 percentage points on AIME, alongside improvements in most code-generation comparisons. Further analysis shows that RECAL reduces residual damage at heavily affected tokens and establishes performance advantages that persist through recovery. These results demonstrate the value of recovery-aware calibration for improving on-policy distillation recovery of pruned reasoning models.
Problem

Research questions and friction points this paper is trying to address.

Structured Pruning
On-Policy Distillation
Reasoning Language Models
Model Recovery
Calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structured Pruning
On-Policy Distillation
Recovery-Aware Calibration
Forward KL Divergence
Reasoning Language Models
🔎 Similar Papers
No similar papers found.
Houcheng Jiang
Houcheng Jiang
University of Science and Technology of China
Model editingLLMs
M
Mao Zheng
Foundation Model Department, Tencent
M
Mingyang Song
Foundation Model Department, Tencent
Q
Qiyong Zhong
Foundation Model Department, Tencent
Jie Sun
Jie Sun
University of Science and Technology of China
T
Tianyu Zhang
University of Science and Technology of China
Junfeng Fang
Junfeng Fang
National University of Singapore
Model EditingAI SafetyLLM ExplainabilityAI4Science