Activation-Conditioned Self-Distillation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that privileged information struggles to provide effective token-level supervision during long-context reasoning. To overcome this, we propose a novel activation-contrastive steering vector extraction mechanism that constructs steering vectors by contrasting activation discrepancies between correct and incorrect reasoning trajectories. Leveraging a frozen model, our approach generates dense supervision signals for self-distillation without requiring reference texts or teacher parameter updates. This method significantly enhances supervision stability at long-tail positions. Extensive evaluations across five models on mathematical benchmarks demonstrate state-of-the-art average accuracy. Notably, when applied to DeepSeek-R1-0528-Qwen3-8B, our approach achieves 71.9% mathematical accuracy and a 70.9% code pass rate, underscoring its effectiveness in improving long-context reasoning capabilities.
📝 Abstract
On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning. Providing privileged information does not by itself ensure effective token-level supervision throughout long responses. We introduce Activation-Conditioned Self-Distillation (ACSD), which extracts a steering vector by contrasting activations of self-generated trajectories that reach verified correct answers within a generation budget with those of all remaining trajectories. A frozen copy of the base model applies this vector at each prediction position, and the student learns from its next-token distributions on student-generated prefixes. Outcome verification is used for direction construction and calibration; distillation requires neither problem-specific reference text nor teacher parameter updates. The distilled student is used alone at inference. On each of five models, ACSD achieves the highest mean accuracy over four mathematical benchmarks among the evaluated methods. On DeepSeek-R1-0528-Qwen3-8B, mean mathematical accuracy reaches 71.9\% and LiveCodeBench v6 pass@12 reaches 70.9\%, compared with 69.0\% and 66.3\% for the reference-conditioned OPSD baseline. Contrasts among correct trajectories also support distillation, and extracted directions can be reused across mathematical training datasets. On fixed student trajectories, ACSD maintains more stable late-position logit-update magnitudes than OPSD.
Problem

Research questions and friction points this paper is trying to address.

self-distillation
token-level supervision
reasoning
on-policy
long responses
Innovation

Methods, ideas, or system contributions that make the work stand out.

Activation-Conditioned Self-Distillation
Steering Vector
On-policy Self-Distillation
Outcome Verification
Representation Engineering
🔎 Similar Papers
No similar papers found.