How Much Can Language Models Gain from Test-Time Computation?

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of unified cross-domain benchmarks and the neglect of selection costs in existing test-time scaling evaluations, which hinder the quantification of genuine reasoning gains. To this end, we propose SELF-POT, a unified evaluation framework that pioneers incorporating model invocation costs measured in USD. It systematically compares direct reasoning with parallel sampling across mathematical, programming, and agent tasks. The framework decouples candidate coverage from final accuracy, encompasses both static and dynamic scenarios, and integrates self-correction, public example selection, and judge techniques with fallback mechanisms under rigorous API cost accounting. Experiments demonstrate that optimizing selection strategies substantially improves accuracy while reducing costs by 12%–49%, revealing the critical role of failure handling in unlocking reasoning potential.
📝 Abstract
How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge selection to the budget. We introduce SELF-POT, a benchmark and evaluation framework that measures the test-time potential of a model across competition mathematics, competitive programming, and agentic workflows. SELF-POT separates candidate coverage from final accuracy on static tasks, tracks correctness transitions under revision, and measures protocol completion alongside task success in agentic environments. Under a unified budget rule, it compares Direct inference with parallel sampling and self-revision under fixed multiples of the Direct budget, and charges every model call, including selection and critique, in dollars. This design supports two kinds of comparison: the gain a model obtains from additional inference, and a lower-cost model with additional inference against a stronger model. Across five low-cost reasoning models on 350 sealed tasks, with Claude Opus 5.5 Direct as the reference, the returns depend on the domain, the selection rule, and failure handling. When we replay the retained programming candidate pools, public-example selection raises correct submissions from 376 to 453 of 500 scheduled cells while saving 12-49% of logical API cost across models, and simply retaining an available candidate when judging fails recovers 61 submissions at unchanged cost. On identical mathematics pools, judging with fallback yields 186 correct submissions versus 182 for voting, while voting saves 12-21% of logical API cost. These controlled replays show how selection and failure handling change the gains realized from the same generated candidates, and they quantify the marginal value of a model judge.
Problem

Research questions and friction points this paper is trying to address.

test-time computation
language models
evaluation benchmark
inference cost
test-time scaling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Computation
Evaluation Framework
Budget-Aware Selection
Agentic Workflows
Self-Revision
🔎 Similar Papers
No similar papers found.
B
Bangji Yang
University of Illinois at Urbana-Champaign
Jingyuan Li
Jingyuan Li
University of Washington
Jiajun Fan
Jiajun Fan
CS Ph.D. of University of Illinois Urbana-Champaign
Reinforcement LearningMachine Learning
Y
Yi Evie Zhang
University of Illinois at Urbana-Champaign
Ruihan Guo
Ruihan Guo
UIUC
Machine LearningDrug Discovery
H
Hongba Ma
Tsinghua University
N
Neil He
University of Illinois at Urbana-Champaign
Chumeng Liang
Chumeng Liang
University of Illinois Urbana-Champaign
Q
Qinglong Zheng
University of Illinois at Urbana-Champaign
Z
Zhanghan Ni
University of Illinois at Urbana-Champaign
Ge Liu
Ge Liu
PhD in CSAIL, MIT; Assistant Professor @ CS, UIUC; Postdoc at IPD, UW
Machine learningcomputational biologyartificial intelligence