ARM: Attention with Routed-Memory for Learnable Sparse Control

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种新的KV缓存结构ARM,通过可学习的稀疏控制方法解决了大语言模型在长上下文推理中的信息丢失问题。
📝 Abstract
Despite advances in long-context inference, large language models (LLMs) remain fundamentally limited by the key-value (KV) caching mechanisms that are necessary for stable computation. Techniques such as selective token eviction and pruning have vastly mitigated these issues, but often discard core information to manage the growing cache. In this paper, we propose Attention with Routed Memory (ARM) a novel KV caching structure that introduces a fully differentiable, fixed-size memory system organized as a hierarchical router. Via a Gumbel-Softmax, ARM learns to select memory slots and perform sigmoid-gated updates that softly combine new and stored information, avoiding hard eviction and reducing information loss. By further training a policy to dynamically select varying amounts of memory at inference, ARM adapts its accesses for both simple contexts and inputs that require deeper reasoning, enabling more scalable and effective retrieval on both short- and long-contexts. Experimental results on standard commonsense and long-context reasoning benchmarks demonstrate that ARM achieves superior performance and efficiency compared to fixed KV-caching approaches, while remaining efficient and scalable in terms of both memory and generation latency.
Problem

Research questions and friction points this paper is trying to address.

large language models
key-value caching
information loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

Attention with Routed Memory (ARM)
Gumbel-Softmax
Sigmoid-gated updates
Hierarchical router
🔎 Similar Papers
2024-02-16Nature CommunicationsCitations: 2