Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the trade-offs between reasoning length, computational cost, and accuracy in large language models, where extended reasoning increases latency and expense, yet the impact of truncating inference remains unclear. The work presents the first systematic evaluation of early-answer strategies under matched reasoning lengths, explicitly distinguishing between completed and incomplete reasoning paths. Controlled experiments are conducted on GPQA Diamond and MMLU-Pro using Qwen3-14B with numerical and concise prompts, alongside gpt-oss-20b and gpt-oss-120b under low- and high-effort training configurations. Results show that concise instructions within 512 tokens improve MMLU-Pro accuracy by 3.8 percentage points. Moreover, low-effort settings that complete reasoning before the cutoff significantly outperform high-effort but incomplete attempts, although no consistent gains are observed in probability mass allocation.
📝 Abstract
Large language models often reason at length before answering, increasing cost and latency. Prompts and trained settings can shorten this reasoning, but a shorter trace may only show that the model stopped sooner. Here, we evaluate paired runs of the same question at matched reasoning horizons across 198 GPQA Diamond and 500 MMLU-Pro questions. We test a numeric/concision prompt that announces a token limit for Qwen3-14B and the trained effort settings of gpt-oss-20b and -120b. The Qwen prompt shortens reasoning traces by 12-17%, while accuracy changes at matched token limits are small and mixed. A concise/early-answer instruction raises MMLU-Pro accuracy by 3.8 percentage points at 512 tokens, including +2.7 points when both runs are unfinished. Its gain at 2,048 tokens is uncertain. For gpt-oss, candidate-logit answers from completed low- and medium-effort reasoning are 14.5-26.3 points more accurate than matched-horizon high-effort answers. Most of the 512-token advantage comes from lower effort finishing earlier, while differences among unfinished runs are smaller and mixed. Wrong early answers often concentrate probability on the chosen option, so earlier stopping does not uniformly improve probability quality. In these tests, a tight deadline can favor lower effort or a concise instruction, whereas allowing high effort to finish can recover higher final accuracy. Evaluations should report correct completion before a deadline, the answer obtained when a run is stopped, differences among unfinished runs, and probability assigned to the correct answer separately.
Problem

Research questions and friction points this paper is trying to address.

reasoning efficiency
early answer
large language models
reasoning trace
accuracy under deadline
Innovation

Methods, ideas, or system contributions that make the work stand out.

reasoning efficiency
early answer
matched reasoning horizon
effort setting
token-limited evaluation
🔎 Similar Papers
No similar papers found.
F
Francesca Carlon
1Data Analytics Lab, Vrije Universiteit Brussel, Pleinlaan 5, 1050 Brussels, Belgium; 2imec-SMIT, Vrije Universiteit Brussel, Pleinlaan 9, 1050 Brussels, Belgium
Vincent Ginis
Vincent Ginis
Vrije Universiteit Brussel / Harvard University
Physics | Machine Learning
A
Andres Algaba
1Data Analytics Lab, Vrije Universiteit Brussel, Pleinlaan 5, 1050 Brussels, Belgium; 2imec-SMIT, Vrije Universiteit Brussel, Pleinlaan 9, 1050 Brussels, Belgium