π€ AI Summary
This study addresses the limitation that large language models, during sequential test-time scaling, tend to prematurely fall into "answer attractors," causing exploration stagnation and constraining long-horizon reasoning performance. To tackle this issue, we present the first quantitative analysis of the attractor phenomenon and propose a model mixing intervention strategy based on recursive self-aggregation. This approach disrupts local optima entrapment to enhance exploratory capacity, while systematically comparing sequential and parallel scaling trajectories. Experimental results demonstrate that the proposed method reduces the attractor hit rate by 21.2% and improves accuracy by at least 2.2%, significantly surpassing the coverage of compute-matched parallel baselines. These findings establish a new paradigm for long-horizon reasoning scaling in large language models.
π Abstract
Test-time scaling improves language model reasoning by spending additional compute at inference. However, both classes of existing methods often fail to continue improving over long timescales. Parallel methods repeatedly sample independent answers from the model, scaling poorly on problems the model is unlikely to solve in a single attempt. In contrast, sequential methods build on previous answers to access new ideas, yet so far have not been shown to reach answers beyond those found by parallel scaling. First, we show that sequential scaling often stops improving because it becomes prematurely trapped in an attractor: a set of answers that prevents exploration of different answers once entered. Across 27 combinations of scaling methods, models, and benchmarks, we find that 53.8% of sequential scaling trajectories enter an attractor within four iterations. Second, we show that a simple model-mixing intervention helps escape attractors. This reduces the attractor hit rate by 21.2 percentage points on average, expands solution coverage beyond a compute-matched parallel baseline, and improves accuracy of recursive self-aggregation by at least 2.2 percentage points. Our results motivate refocusing long-horizon test-time scaling from parallel methods to sequential methods that improve previous answers.