Decoding Looped Transformers Better for (Almost) Free

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of recurrent Transformers, which discard early intermediate states across iterations, thereby constraining decoding quality and introducing computational redundancy. To overcome this, we propose LoopCD, a training-free contrastive decoding framework that pioneers the exploitation of intrinsic alignment signals within the recurrence. Specifically, LoopCD performs contrastive decoding between weak and strong prediction pairs in both the logit and hidden state spaces to optimize token selection. Empirical evaluations on the AIME and HumanEval benchmarks demonstrate that our method significantly improves accuracy while enabling the number of recurrence steps to be halved. Consequently, LoopCD reduces inference FLOPs by 22.5% to 48.2%, achieving simultaneous gains in both performance and computational efficiency.
📝 Abstract
Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. We introduce LoopCD, a training-free contrastive decoding framework that guides token selection by contrasting the final prediction with an earlier recurrent pass, operating either in logit space with one extra output pass (LoopCD-Logits) or in hidden-state space with zero output overhead (LoopCD-Hidden). Across four looped Transformer families, LoopCD delivers substantial, consistent gains at full recurrent depth: LoopCD-Logits raises Ouro-2.6B-Thinking's AIME 2024 pass@1 from 61.88% to 73.33%, while LoopCD-Hidden lifts Huginn's HumanEval pass@1 from 22.56% to 31.71%. Crucially, these performance gains enable halving the number of recurrent loops while still matching or exceeding full-depth unguided baselines, reducing forward FLOPs by 22.5% to 48.2%. By transforming intermediate recurrent states into effective guidance signals, LoopCD achieves superior decoding quality while substantially reducing inference compute.
Problem

Research questions and friction points this paper is trying to address.

Looped Transformers
Contrastive Decoding
Intermediate Representations
Inference Efficiency
Parameter Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Looped Transformers
Contrastive Decoding
Training-free
Parameter Efficiency
Inference Compute Reduction
🔎 Similar Papers
No similar papers found.