🤖 AI Summary
This study addresses the prohibitive computational overhead in recurrent language models caused by redundant computations during training, decoding, and reinforcement learning. To mitigate this, the proposed method leverages fixed-point properties to enable truncated backpropagation through time and KV cache sharing. It further introduces a mechanism for learning deep priors from predictive feedback with entropy regularization, combined with orthogonal injection to eliminate interference among state components. Additionally, knowledge distillation is integrated with reinforcement learning gradient optimization to enhance overall efficiency. Experimental results demonstrate that the approach accelerates prefilling by 1.79× and reinforcement learning updates by 2×. Notably, with 1.6 billion parameters, the model achieves performance comparable to full-cache baselines while utilizing only one-third of the KV cache size.
📝 Abstract
Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables truncated backpropagation in training; terminal key-value (KV) sharing for decoding with almost no loss in accuracy; a distilled student that prefills up to 1.79x faster; and RL updates that compute gradients from saved rollout states, 2x faster than backpropagating through the replayed trajectory. We therefore improve the two components of training that shape these fixed points: the depth prior and input injection. Fixed-depth training breaks KV sharing, and Huginn's broad depth prior supports sharing but dilutes supervision at the target depth more than sharing requires; we learn the prior from prediction feedback, with an entropy term that keeps it broad. Existing injection schemes let the state's component along the input amplify or cancel the injection; we remove this component with orthogonal injection. From 100M to 1.6B parameters, the learned prior and orthogonal injection lower perplexity at every scale relative to Huginn's prior and existing injection schemes, respectively. At 1.6B, the learned prior with a 3x smaller KV cache matches the downstream average of fixed-depth training with the full cache.