Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
该研究针对递归语言模型中全局注意力重复计算问题,提出WISE方法,在早期使用全局注意力发现相关上下文,并在后期重用已发现的支持集以提高效率。
📝 Abstract
Recurrent language models repeatedly apply shared network blocks to refine latent representations, but standard inference recomputes global attention at every recurrent step. We study attention dynamics across recurrent depth and find that attention support and distributions stabilize substantially earlier than hidden states and attention outputs. This suggests a two-stage structure: early steps discover a sparse working set of relevant context, while later steps refine representations over largely the same routing support. Motivated by this structure, we introduce WISE (Working-set Inference with Support Exploitation), a training-free method that uses unrestricted global attention during early recurrence and later reuses directly discovered block-structured support while keeping recurrent depth and within-support attention computation dynamic. Controlled interventions show that recurrent discovery is important and that support-only reuse better preserves model behavior than more restrictive attention-reuse alternatives. Across multi-hop QA benchmarks, WISE largely preserves full-attention performance, while context scaling reveals increasingly sparse working sets and greater efficiency gains. Quality is largely preserved through 2K context, with a measurable loss at 4K. An optimized sparse-attention implementation achieves up to a 1.76x attention speedup over native FlashAttention at 4K and a 1.36x speedup for the full 32-step attention trajectory. Our code is available at https://github.com/tbn5pj/WISE_code.
Problem

Research questions and friction points this paper is trying to address.

Attention Routing
Recurrent Language Models
Working-Set Inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recurrent Language Models
Attention Routing
Sparse Working Set
Support Reuse
Efficiency Gains
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Ke Wan
Department of Computer Science, University of Virginia
Chen Chen
Chen Chen
University of Virginia
Data MiningMachine LearningComputational Epidemiology