🤖 AI Summary
This study addresses the prohibitive computational overhead of redundant historical context processing and the rapid growth of KV cache memory with increasing sequence length in long-context language models. To this end, it proposes a shared global state architecture that introduces a novel causal encoder-decoder framework. During the encoding phase, multi-stage retrieval and hierarchical query update strategies aggregate long-range historical information into a fixed-window shared state, which the decoder directly reuses. This mechanism eliminates the need for redundant, layer-wise reconstruction of historical representations. Consequently, the proposed approach significantly reduces both computational costs and cache overhead while preserving model performance and its capacity to effectively leverage long-range dependencies.
📝 Abstract
Efficient language models must reduce not only the cost of individual accesses to past context but also the overhead of repeatedly selecting and processing historical information across layers. We introduce the Global State Model (GSM), a causal encoder--decoder architecture that concentrates the selection and aggregation of long-range information in the encoding stage. Through multiple stages of history retrieval, the encoder progressively incorporates long-range information into representations at recent positions, forming a shared state with a fixed window size. Each decoder layer accesses this same state using queries updated from the preceding layer, preserving computational depth while avoiding repeated construction of historical key--value (KV) representations and long-range indexing. As a result, neither the decoder's per-step attention cost nor its KV cache size grows with the history length. Experiments show that GSM improves computational efficiency and reduces cache overhead while maintaining model performance and the ability to use long-range information, offering a shared-state architecture for efficient language modeling.