Improving Test-Time Scaling with Adaptive Looped Transformers

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of recurrent Transformers in test-time scaling, where fixed-depth architectures lead to computational waste and accuracy bottlenecks. To overcome this, we propose TaH2, a framework that jointly trains an iterative decision-maker for adaptive compute allocation. The architecture introduces a lookahead depth supervision mechanism alongside online label supervision, enabling post-training techniques to selectively increase iterations only on beneficial tokens, thereby breaking the fixed-depth constraint. Evaluated on the AIME benchmark, TaH2 improves the test-time scaling slope by 53% and surpasses baseline peak accuracy by 3.4 percentage points under equivalent computational budgets, significantly optimizing both inference efficiency and performance upper bounds.
📝 Abstract
Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs grow longer remains underexplored. Through post-training looped transformers, we study the accuracy-compute slope, measured as the accuracy gain per doubling of test-time decoding FLOPs. We find that existing looped transformers often yield steeper slopes than their non-looped baseline, yet underperform it at matched compute. While fixed-depth looping spends extra iterations on every token, our analysis shows that many tokens do not benefit from extra iterations. We therefore propose TaH2, which enables the model to focus extra iterations on the tokens that benefit from looping. It jointly post-trains the backbone and an iteration decider through lookahead depth supervision, which uses online labels indicating whether further iteration improves the prediction. TaH2 improves both the efficiency and attainable accuracy of test-time scaling. On challenging AIME benchmarks, TaH2 improves the accuracy-compute slope by 53% (2.74 vs. 1.79) over the non-looped baseline, exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute. As the maximum iteration depth increases, existing looped models largely plateau, while TaH2's gain over the non-looped baseline continues to grow from +2.8 points at depth 2 to +3.9 points at depth 8. Our code is available at https://github.com/thu-nics/TaH.
Problem

Research questions and friction points this paper is trying to address.

Looped Transformers
Test-Time Scaling
Adaptive Computation
Parameter Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Looped Transformers
Test-Time Scaling
Adaptive Computation
Lookahead Depth Supervision
Iteration Decider
🔎 Similar Papers
No similar papers found.