On-Policy Attention Linearization

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance collapse caused by error accumulation in linear attention during long-context inference within hybrid architectures. To mitigate this, we propose OPAL, an online policy distillation method wherein a student model samples its own trajectories and incorporates an online policy mechanism to correct distribution shift. By leveraging dense supervision from a teacher model, OPAL achieves efficient linearized distillation, restoring long-sequence processing capabilities without requiring supervised fine-tuning (SFT) or reinforcement learning with verifiable rewards (RLVR). Trained on merely 3 billion tokens, OPAL recovers 87%–94% of common-sense reasoning and 100% of retrieval performance, while attaining 83%–93% in mathematical reasoning. These results demonstrate that OPAL provides an efficient solution for deploying hybrid Transformer architectures in long-context scenarios.
📝 Abstract
Hybrid transformer architectures that replace most softmax attention layers with linear attention offer transformer-level quality at a fraction of the memory cost. Rather than pretraining such models, a growing body of work distills them from already trained full-attention transformers. However, these distilled models often collapse on long-context retrieval and reasoning tasks, particularly when operating in thinking mode, where the efficiency gains of hybrid architectures matter most. Since linear attention layers must compress context into a fixed-size state, their errors compound over long sequences. As off-policy distillation never teaches the student model to recover from this drift, tasks that necessitate longer sequence lengths become especially challenging. We introduce On-Policy Attention Linearization (OPAL) in which the hybrid attention student samples its own long-context trajectories and receives dense supervision from the frozen full-attention teacher. Applying OPAL to Qwen3-4B and MiMo-7B-RL-0530, we recover $87$--$94\%$ of full-attention performance on commonsense reasoning, $100\%$ on needle-in-a-haystack (NIAH) retrieval, and $83$--$93\%$ on mathematical reasoning with only 3B training tokens. We achieve these results without supervised fine-tuning (SFT) or reinforcement learning with verifiable rewards (RLVR). Compared with the strongest prior linearization method, which recovers $68\%$ of its teacher's retrieval performance and $21.6\%$ absolute average mathematical reasoning accuracy, OPAL fully recovers retrieval and achieves $67.6$--$72.2\%$ on math reasoning.
Problem

Research questions and friction points this paper is trying to address.

linear attention
knowledge distillation
long-context retrieval
error accumulation
hybrid transformer
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Linear Attention
Hybrid Transformer
Attention Linearization
Long-context Reasoning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Arian Raje
Carnegie Mellon University
A
Anupam Nayak
Carnegie Mellon University
A
Anthony Fei
Cornell University
A
Akaash Parthasarathy
Carnegie Mellon University
M
Mohamed Abdelfattah
Cornell University
Gauri Joshi
Gauri Joshi
Carnegie Mellon University
applied probabilitymachine learningoptimizationinformation theory