Enabling Dynamic Computation in Looped LMs

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that recurrent language models constrained by KV caches cannot perform dynamic computation, and that existing early-exit mechanisms enforce a static processing depth for all tokens. To overcome this, we propose a "best-available" KV cache strategy alongside an improved early-exit prior objective, introducing the first KV cache scheme compatible with dynamic computation. Coupled with a parameter-efficient architecture, this approach enables difficulty-based heterogeneous depth exiting. Our contributions break the performance–computation-depth trade-off bottleneck, reducing both FLOPs and KV memory consumption by 30% while preserving full-depth performance, thereby significantly enhancing model efficiency and flexibility.
📝 Abstract
Looped LMs are parameter efficient and promise dynamic computation (saving memory and FLOPs on easy tokens). However, state-of-the-art open Looped LMs trained with this dynamic computation capability (Ouro models) do not realize it in practice as each loop iteration (depth) requires its own level of KV-cache, necessitating all loop computations. Moreover, Ouro's early-exit prior is enforced on each token equally, which results in static lower-depth like processing of all tokens regardless of difficulty. In this work, we propose a simple "best-available" KV caching strategy that works out-of-the-box, creating a new frontier in the performance vs depth space. Our approach enables up to 30% reduction in FLOPs and KV memory while retaining full-depth performance, showing the true flexibility of Looped LMs. Furthermore, training looped LMs with awareness about this KV caching strategy improves performance and efficiency. Finally, we apply a small but effective fix to the early-exit prior enforcement objective that makes tokens exit at truly heterogeneous depths based on effort. Our findings are validated on Ouro models as well as smaller looped LMs pre-trained from scratch.
Problem

Research questions and friction points this paper is trying to address.

Looped Language Models
Dynamic Computation
KV-cache
Early-exit
Parameter Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Looped Language Models
Dynamic Computation
KV Caching Strategy
Early-Exit Prior
Parameter Efficiency
🔎 Similar Papers
No similar papers found.