Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing online self-distillation methods for reasoning models, which supervise only token distributions while neglecting attention positions. To this end, we propose OPASD, introducing a pioneering conditionalized attention distillation mechanism that projects and normalizes teacher attention onto student-visible positions to provide precise contextual guidance. By integrating attention projection and renormalization into an online policy self-distillation framework, OPASD effectively complements the shortcomings of token-level supervision. Experimental results demonstrate that OPASD improves accuracy by 4.98%–8.4%, reduces generated tokens by 73.9%, and decreases computational cost by 72.6%, achieving significant synergistic optimization of performance and efficiency.
📝 Abstract
On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding context. We introduce On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation. Because the privileged teacher can attend to verified solution tokens unavailable to the student, OPASD projects teacher attention onto student-visible positions and renormalizes the resulting distribution before alignment. Across three model sizes and four competition-level mathematics benchmarks, OPASD consistently outperforms token-only OPSD, improving average accuracy by 4.98 to 8.40 percentage points. OPASD also avoids the response-length inflation and performance degradation observed with token-only distillation, reducing generated rollout tokens by 73.9% and estimated model compute by 72.6% while training 1.53x faster. These results show that solution-conditioned attention provides a complementary supervision signal that makes on-policy self-distillation more accurate, stable, and compute-efficient.
Problem

Research questions and friction points this paper is trying to address.

On-policy self-distillation
Reasoning models
Attention distillation
Token-level supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Attention Self-Distillation
On-Policy Distillation
Reasoning Models
Attention Projection
Compute Efficiency
🔎 Similar Papers
No similar papers found.