Pooling Helps, Learned Weighting Hurts In-Context: Decomposing Group Attention

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear contributions of individual components within grouped attention mechanisms to multivariate forecasting and in-context learning. To this end, it presents the first decoupled analysis of the internal workings of grouped attention by editing attention matrices during the inference phase to isolate V/O projection pathways from Q/K weight pathways, combined with hierarchical ablation experiments incorporating a uniform pooling strategy. The findings reveal the counterintuitive phenomenon that uniform pooling outperforms learned weighting. Specifically, uniform pooling proves effective across most configurations, whereas learned weighting degrades in-context learning performance. Furthermore, applying uniformization exclusively at the first layer suffices to comprehensively enhance overall model performance.
📝 Abstract
Group attention, introduced by the time series forecasting model Chronos-2, attends over the variates of a group at a fixed patch index and serves both multivariate (MV) and in-context learning (ICL) forecasting. Rather than evaluating this cross-variate attention design as a whole, we ask which part of the mechanism earns the benefit and probe its applicability to both MV and ICL regimes. By editing the attention matrix $α$ at inference we separate the two pathways a head comprises: V/O, which projects a weighted summary of the group, and Q/K, which decides the weights. Uniform pooling (V/O without any Q/K weighting) is positive on 18 of our 20 sensor-network configurations, while the learned weighting (Q/K) splits by group type: its contribution is positive or negligible for MV, but materially degrades 8 of the 10 sensor-network ICL configurations, leaving 4 of them worse than univariate inference. By isolating the impact of different layers, we find that uniforming $α$ in the first block alone improves every ICL configuration we test.
Problem

Research questions and friction points this paper is trying to address.

Group Attention
In-Context Learning
Multivariate Forecasting
Time Series
Attention Decomposition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Group Attention
In-Context Learning
Attention Decomposition
Time Series Forecasting
Uniform Pooling
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
M
Michael Fore
Amazon
J
James Mason Inder
Amazon
M
Mrishika Nair
Amazon
Praneetha Vaddamanu
Praneetha Vaddamanu
Applied Scientist, Microsoft Turing
Computer ScienceArtificial Intelligence
S
Sharlina Keshava
Amazon