On the Trade-off Between Information Loss and Generalization in Sparse Attention

πŸ“… 2026-10-03
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the theoretical gap in sparse attention mechanisms and the unclear relationship between information fidelity and generalization by systematically establishing their theoretical connection for the first time. Methodologically, it quantifies approximation error via Jensen-Shannon divergence and leverages concentration analysis of order statistics to derive generalization bounds based on Rademacher complexity and mutual information. The research yields closed-form error expressions and sparsity-dependent generalization upper bounds. These results reveal a fundamental trade-off wherein sparsity reduces hypothesis class complexity while introducing information loss, thereby elucidating the theoretical limits of sparsification mechanisms in Transformers.
πŸ“ Abstract
To mitigate the quadratic complexity bottleneck of the Transformer, sparse attention has emerged as a pivotal technology. Despite the extensive empirical success of sparse Transformers, the theoretical understanding of sparse attention remains fragmented. In particular, two fundamental questions remain unclear: (1) How does sparsification affect the information fidelity of attention mechanisms? (2) How does this information loss interact with the generalization behavior of the model? To bridge this gap, this paper proposes a systematic analysis of the Jensen-Shannon (JS) divergence and of the generalization gap of sparse attention mechanisms. Specifically, we first characterize the approximation error via the JS divergence. Through an order-statistics-based concentration analysis of the truncation mass alpha --- where the attention scores are assumed to be independent and identically distributed sub-Gaussian random variables with parameter sigma --- the JS divergence between the full attention distribution and the sparse attention distribution is shown to admit the closed form log 2 + ((1 - alpha)/2) log(1 - alpha) - ((2 - alpha)/2) log(2 - alpha). Subsequently, we derive a generalization bound through Rademacher complexity, quantified by O(gamma * sqrt(M/n) * (sqrt(log(3eL/M)) + sqrt(pi)/2)). Furthermore, building on a mutual-information-based generalization bound together with an entropy and covering-number analysis of the sparse hypothesis class, we obtain the sparsity-dependent generalization bound O(sqrt((M/(2n)) * (log(eL/M) + log(1 + 2/epsilon)))). Our analysis shows that sparsity reduces the complexity of the considered hypothesis class while introducing approximation error that can be quantified by the JS divergence. These findings provide a theoretical characterization of the trade-off between information fidelity and generalization in sparse Transformer architectures.
Problem

Research questions and friction points this paper is trying to address.

Sparse Attention
Information Loss
Generalization
Transformer
Trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Attention
Jensen-Shannon Divergence
Generalization Bound
Rademacher Complexity
Information Loss
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Z
Zhongqi Fan
Beijing Normal–Hong Kong Baptist University
Zheng Tan
Zheng Tan
University of California, Los Angeles
Numerical AnalysisImage Processing