🤖 AI Summary
This study addresses the limitation of existing sequence-level power sampling, which employs a fixed exponent to sharpen distributions while disregarding variations in query difficulty and model capability. We propose an adaptive power sampling method grounded in self-reward gap theory, achieving the first training-free, query-level dynamic sharpening. Specifically, during test time, the optimal sharpening exponent is adaptively determined for each query based on answer consistency and self-reward relationships. This approach consistently outperforms fixed-exponent baselines on reasoning benchmarks such as MATH500, significantly enhancing the test-time reasoning performance of large language models with zero additional training cost.
📝 Abstract
Sequence-level power sampling has recently emerged as a training-free approach to reasoning by sampling from a sharpened output distribution of a base large language model (LLM). Nevertheless, existing methods typically sharpen the base model distribution uniformly across queries, overlooking variations in query difficulty and in how well the base model already handles each query. The goal of this work is to equip power sampling with query adaptivity. Theoretically, we show that the benefits of further sharpening are determined by the self-reward gap between correct and incorrect responses. Based on this insight, we propose \emph{Adaptive Power Sampling} (APS), which adjusts the sharpening exponent on a per-query basis at test time using the relationship between answer agreement and the model's self-reward. Experiments across diverse reasoning tasks, including MATH500, HumanEval, and GPQA, show that APS consistently outperforms power sampling with a fixed sharpening exponent, without additional training.