Hybrid Latent Attention for Looped Language Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high inference latency and limited concurrency caused by KV cache bloat during the decoding phase of recurrent language models. To this end, we propose a hybrid latent attention mechanism that employs a sliding window to retain recent exact KV pairs while compressing historical tokens into compact latent vectors for direct query access, eliminating the need to reconstruct original KV states. During training, pretrained weights are kept frozen, and upsampling techniques are leveraged to fine-tune only newly introduced parameters, thereby preserving the original attention behavior. Experimental results demonstrate that the proposed method achieves over 10.7× cache compression while retaining more than 97% accuracy. Furthermore, it increases concurrency by 4× to 8.8× and boosts throughput by up to 7.4×.
📝 Abstract
Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key-value (KV) cache by T. The larger cache limits how many sequences a GPU can decode at once and slows each decoding step, which reads the whole cache. We propose Hybrid Latent Attention (HLA), which keeps exact keys and values within a sliding window of W recent tokens and stores each older token as a compact latent that the query of each loop reads directly, without reconstructing keys and values. We uptrain HLA on Ouro looped models (T=4) with 1.4B and 2.6B parameters, keeping the pretrained weights frozen and training only the added parameters to reproduce the original attention. The cache shrinks by 10.7x per token, fitting 4.0-8.8x as many concurrent sequences per GPU, and decoding throughput improves by 2.5x at 1K-token contexts and by up to 7.4x at 16K. HLA retains over 97% of the original accuracy on math, knowledge and reasoning benchmarks, and 96-100% on long-context retrieval up to 16K tokens. After supervised fine-tuning, it performs on par with the fine-tuned original model on competition-level math.
Problem

Research questions and friction points this paper is trying to address.

looped language models
KV cache
decoding throughput
memory efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid Latent Attention
Looped Language Models
KV Cache Compression
Sliding Window
Decoding Throughput