๐ค AI Summary
Existing test-time scaling methods often discard valid prefixes due to insufficient fine-grained error localization, leading to wasted computation. This work proposes the TTEL algorithm, which introduces the first token-level error detection mechanism by comparing conditional probabilities against a null-context baseline. By truncating erroneous trajectories and regenerating only the faulty segments while preserving correct prefixes, TTEL substantially improves the Pareto frontier of inference efficiency and performance. On LiveCodeBench, it achieves a 71.0% pass@64 rate using approximately half the number of generated tokens compared to prior approaches, and consistently outperforms existing test-time scaling methods across multiple mathematical reasoning benchmarks.
๐ Abstract
Scaling inference-time computation has emerged as a reliable method to improve the performance of large language models on complex reasoning and programming tasks. However, standard approaches such as independent sampling and sequential multi-turn refinement operate without token-level credit assignment, resulting in computational inefficiency, since valid reasoning prefixes are frequently discarded. In this work, we introduce Test-Time Scaling via Error Localization (TTEL), an inference-time algorithm that utilizes fixed or environment feedback to perform token-level error localization. By comparing conditional probabilities under informed feedback against a null-context baseline, TTEL isolates the step at which an error occurred. The algorithm then truncates the trajectory and branches a new generation, maximally reusing the valid prefix. Extensive evaluations demonstrate that TTEL establishes strictly dominating Pareto frontiers across sequential reasoning domains, measured by pass-at-k vs. generated-token cost. With Qwen3-8B on LiveCodeBench, TTEL attains a pass@64 of 71.0% while generating approximately half as many tokens as independent sampling (360.4k vs. 735.0k). Generalizing to math benchmarks AIME-2025 and HMMT-2025, TTEL cleanly outperforms competing test-time baselines across both Qwen3-8B and Qwen3-4B-Thinking-2507.