🤖 AI Summary
This study addresses the prohibitive computational overhead caused by attention parameter scaling in long-context scenarios by proposing the NAMOH architecture. This method introduces a natively sparse attention mechanism that achieves hybrid head sparse selection via K-of-H head activation, jointly determining active parameters and available context without requiring full-history scanning. Furthermore, it incorporates subsequence causal attention and head-relative rotary position encoding. Experimental results demonstrate that, under equivalent parameter budgets, NAMOH outperforms dense models while significantly enhancing long-context inference efficiency. The proposed architecture is also compatible with existing mechanisms such as Grouped Query Attention (GQA), thereby achieving synergistic and efficient scaling of both parameters and context.
📝 Abstract
Scaling attention parameters can improve language model quality, but retaining full token histories makes additional heads costly at long contexts. Furthermore, since attention retrieves and combines contextual information, parameter scaling should also support longer contexts. We therefore ask whether attention parameter scaling can directly enable efficient and effective context scaling. We introduce NAMOH, an architecture-native sparse attention mechanism that activates $K$ of $H$ heads per token. Each head retains only its assigned tokens and performs causal attention within this subsequence. Head selection thus jointly determines active parameters and available context without scanning the full history. Under balanced assignments, increasing $H$ at fixed $K$ shortens head histories and reduces per-token key-value (KV) access without increasing total KV storage. We further support head-relative rotary position embeddings to shorten positional spans within routed subsequences, aiming to mitigate position-induced attention noise. Experiments show that NAMOH can outperform fully activated models with the same total parameters, while enabling more efficient long-context inference than smaller dense models with matched active parameter counts. It remains compatible with GQA and existing sparse attention mechanisms. We hope this work offers a new path for scaling attention, with parameter scaling directly enabling context scaling.