Scheduling Recursive Reasoning in Looped Transformers

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Recurrent Transformers struggle with test-time compute scaling due to fixed update scales, often resulting in overly conservative or aggressive inference. This work proposes TAPS, a scheduler that treats the update scale as an independent control axis and adaptively modulates recursive step sizes online via a progress-fluctuation decomposition to optimize inference trajectories. Enabling dynamic scheduling without retraining, TAPS integrates loss sensitivity analysis with an online step-size adjustment algorithm to substantially improve efficiency. Experimental results demonstrate that TAPS effectively enhances terminal accuracy, achieving up to 1.56× speedup while matching baseline precision, and exhibits strong generalization capabilities across diverse settings.
📝 Abstract
Recurrent reasoning models have attracted growing attention for scaling test-time computation, typically by iteratively refining latent states with shared parameters. However, these models apply each learned update with a fixed unit scale, which can be conservative when updates make persistent progress and overly aggressive when they fluctuate, limiting the benefit of additional loops. To understand how the scale should vary along the trajectory, we first analyze the sensitivity of terminal loss to recurrent update scale. We show that its temporal average admits an exact decomposition into persistent-progress and centered-fluctuation contributions. Based on this, we introduce the Trajectory Adaptive Progress-Fluctuation Scheduler (TAPS), which tracks their balance across recurrent updates and adapts the step size online. Theoretically, we establish sufficient conditions under which TAPS reduces expected terminal loss and reaches a target quality in fewer recurrent loops. Empirically, we show that TAPS improves terminal accuracy across structured reasoning tasks without retraining. By further incorporating the progress-fluctuation principle into training, TAPS yields additional accuracy gains with up to 1.56 times wall-clock speedup at matched baseline accuracy. The broad applicability of TAPS is supported by its effectiveness across diverse recurrent architectures and inference strategies. Together, these results establish update scale as complementary control axis of recurrent inference alongside architecture and depth.
Problem

Research questions and friction points this paper is trying to address.

Looped Transformers
Recurrent Reasoning
Test-time Compute Scaling
Update Scale Scheduling
Step Size Adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Looped Transformers
Recursive Reasoning
Adaptive Step Size
Progress-Fluctuation Decomposition
Test-Time Compute
🔎 Similar Papers
No similar papers found.