🤖 AI Summary
This study addresses the limitation of existing self-distillation methods, which are confined to the training phase and inapplicable during inference. To overcome this, we propose a test-time self-distillation framework that introduces self-distillation principles into the inference stage for the first time. Specifically, our method employs counterfactual context templates in place of expert demonstrations to extract guiding signals and constructs rewards based on log-odds ratios. By approximating the target distribution through Gibbs reweighting and integrating beam search for global trajectory optimization, the approach enhances decoding quality without requiring parameter updates. Experimental evaluations across mathematical reasoning, code generation, and scientific question-answering benchmarks demonstrate that the proposed framework significantly outperforms standard sampling, low-temperature decoding, and conventional beam search baselines.
📝 Abstract
Self-Distillation Fine-Tuning (SDFT) enables a language model to act as its own teacher: by conditioning on a demonstration, the model produces an implicit reward via pointwise mutual information, which guides on-policy learning without external supervision. However, SDFT operates at training time: it requires gradient updates and access to expert demonstrations, making it inapplicable at inference. We propose test-time self-distillation, a decoding-time method that extracts a steering signal from the self-distillation framework without any parameter updates, reward models, or training data. Our key insight is that counterfactual contexts, i.e. fixed textual templates that hypothetically prime the model for excellent versus poor reasoning, can substitute for the demonstration. The log-odds ratio of a candidate answer under these two counterfactual conditions defines a new reward signal. We derive the optimal KL-regularized policy under this reward, which takes the form of a Gibbs reweighting of the base distribution. Crucially, this reweighting is global: it cannot be decomposed into independent per-token operations without ignoring future trajectory quality. We therefore approximate the target distribution via beam search. Experiments on mathematical reasoning (MATH500), code generation (HumanEval), and graduate-level science QA (GPQA) across multiple model scales show that test-time self-distillation improves over standard sampling, low temperature, beam search and power sampling baselines on average, demonstrating that the self-distillation principle can be operationalized at inference time.