MassAlloc Attention: Let Attention Allocate Its Own Compute

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses computational waste from low-contribution operations and redundancy in dense kernel execution within full attention mechanisms by proposing MALA, a fused attention primitive. MALA pioneers the use of an online Softmax normalizer to dynamically allocate post-score computation, pruning low-value operations while preserving complete causal access. This unifies training and inference tolerances, achieving distribution-adaptive resource optimization. Furthermore, it enhances execution efficiency through fused attention kernels, nested retention support derivation, and standard attention state reuse. Experiments demonstrate that at 128K context length, MALA reduces training latency by 2.2× to 3.0× and accelerates decoding by 1.6×, matching full attention performance while significantly reducing FLOPs.
📝 Abstract
FullAttn often assigns negligible normalized mass to much of the causal score space, yet dense kernels execute the complete post-score path after forming each QK tile. We introduce MALA, a fused attention primitive that preserves score access to every legal causal interaction and uses normalized contribution to allocate post-score computation. Forward uses its evolving online-softmax normalizer, while backward reuses the finalized normalizer to derive nested retained support using only standard attention state. A common tolerance governs training and inference, allowing for adaptive retention of the work. MALA reduces low-contribution post-score computation. A matched-work study at 8K isolates the benefit of distribution-adaptive allocation: under exactly matched total post-score work, MALA approaches a per-instance reference-mass oracle, with mean omitted mass of 0.0188% versus 0.0182%. Across context lengths from 1K to 32K tokens, the same tolerance maintains low output and gradient errors relative to the reference. Across a broader controlled associative-recall comparison, MALA closely tracks FullAttn as context grows, reaching 89.67% accuracy at 8K compared with 89.97% for FullAttn. In an attention-operator benchmark at 128K tokens with tensor parallelism, MALA reduces forward and backward latency during training by 2.2x and 3.0x and decoding latency during inference by 1.6x relative to FullAttn. Across scaling-law training from 0.6B to 14B parameters, MALA closely tracks FullAttn in perplexity while reducing total training FLOPs. The resulting 14B models and 32B models from separate continued training achieve comparable knowledge, reasoning, and long-context retrieval scores to FullAttn. These results indicate that allocating post-score computation according to normalized attention contributions can retain the evaluated capabilities of FullAttn while reducing attention computation.
Problem

Research questions and friction points this paper is trying to address.

Attention Mechanism
Compute Allocation
Sparse Attention
Efficient Transformers
Causal Masking
Innovation

Methods, ideas, or system contributions that make the work stand out.

MassAlloc Attention
Fused attention primitive
Adaptive compute allocation
Online-softmax normalizer
Distribution-adaptive allocation
J
Jingze Shi
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China; Beijing Academy of Artificial Intelligence, Beijing, China
Zhangyang Peng
Zhangyang Peng
Master Student, Hangzhou Dianzi University
Vector DatabaseApproximate Nearest Neighbor SearchCommunity Search
X
Xianduo Li
Beijing Academy of Artificial Intelligence, Beijing, China
Yanlin Qi
Yanlin Qi
University of California, Davis
Traffic PredictionsUrban ComputingData MiningGeoAI
X
Xiaotian Lin
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
H
Haoxian Chen
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China; Beijing Academy of Artificial Intelligence, Beijing, China
L
Liangdong Wang
Beijing Academy of Artificial Intelligence, Beijing, China
Guang Liu
Guang Liu
BAAI
AI,LLMData
Yuyu Luo
Yuyu Luo
Assistant Professor, HKUST(GZ) / HKUST
Data AgentsLLM AgentsDatabaseText-to-SQLData-centric AI