Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the exploration-exploitation trade-off in inference-time sampling and the prohibitive costs and limited generalization associated with reinforcement learning-based post-training. To this end, we propose Parallel Power Tempering, a method that optimizes multi-chain replica interactions and exchange strategies to effectively balance diverse exploration with high-probability exploitation. Furthermore, it incorporates a truncation bias correction mechanism to achieve inference-time compute scaling without parameter updates. Experimental results demonstrate that our approach substantially improves the sampling quality of smaller models, surpassing the performance of reinforcement learning post-trained counterparts while achieving results comparable to those of frontier large language models.
📝 Abstract
Power-sharpened sampling is an inference-time alternative to reinforcement-learning (RL) post-training for enhancing reasoning in large language models (LLMs). High-probability sequences are amplified under the base model without parameter updates or external rewards, avoiding the costly optimization and jagged generalization of RL. However, this approach faces a fundamental exploration--exploitation trade-off, as % strong sharpening restricts exploration, trapping samplers in plausible but incorrect reasoning trajectories, whereas weak sharpening leaves the answer distribution diffuse. To resolve this trade-off, we introduce \textbf{Parallel Power Tempering (PPT)}, instantiating power-sharpened LLM sampling via parallel tempering. Running multiple \emph{interacting} replicas in parallel at different sharpening levels allows lower-power replicas to explore diverse reasoning trajectories and higher-power chains to further exploit higher-likelihood responses favored by the sharpened target. Specifically, we tailor \method{} to inference-time sampling by mitigating a truncation bias, identified in prior power samplers, and investigate effective swap strategies under finite memory and compute budgets. Extensive experimentation shows that \method{} substantially improves single-chain power-sharpened sampling and outperforms RL-post-trained models, producing higher-quality reasoning traces and even achieving performance comparable to frontier models.
Problem

Research questions and friction points this paper is trying to address.

power-sharpened sampling
exploration-exploitation trade-off
reasoning
large language models
inference-time sampling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parallel Power Tempering
Power-sharpened sampling
Inference-time reasoning
Exploration-exploitation trade-off
Truncation bias mitigation
P
Panagiotis Theodoropoulos
Georgia Institute of Technology
N
Nan Jiang
University of Texas at El Paso
X
Xintong Duan
ML Research, Morgan Stanley
Ali Hasan
Ali Hasan
Duke University
Y
Yuriy Nevmyvaka
ML Research, Morgan Stanley
E
Evangelos A. Theodorou
Georgia Institute of Technology
Wei Deng
Wei Deng
ML Researcher, Morgan Stanley
Monte CarloDiffusion Models