Emergent Structure in the Marginal Attention Space of Language Models

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inadequate characterization of attention mechanisms in independently trained large language models by constructing a joint Token-Head marginal attention space. This formulation reveals intrinsic textual signals along the token axis and model-specific signatures along the head axis, while establishing their theoretical connection to input-output Jacobian matrices. Building upon post-softmax attention marginalization and Jacobian-based statistical analysis, we propose a training-free strategy for per-head KV cache budget allocation and eviction. Extensive evaluations on standard benchmarks demonstrate that this approach achieves performance comparable to methods requiring recomputation or fine-tuning. The source code has been made publicly available.
📝 Abstract
While representation similarity across independently trained language models is well-documented, how internal mechanics such as attention behave across models remains far less characterized. Inspired by this gap, we examine the structure of post-softmax attention weights by marginalizing over query positions, mapping them into a joint token-head "marginal attention space". Evaluating across 60+ diverse LLMs, we find that different properties emerge when reducing this space along its token and head axes. When reduced token-wise, marginal attention yields a text-intrinsic signal robustly conserved across models. To explain this property, we empirically connect marginal attention to the input-output Jacobian of the network, and prove theoretically that under a smoothness assumption, models with similar next-token distributions are guaranteed to have similar input-output Jacobian statistics. When reduced head-wise, it forms a model-private signature conserved across documents. Practically, this provides a natural way to estimate a per-head budget for key-value (KV) cache eviction, effectively decoupling model-specific budget allocation from text-intrinsic token scoring. On standard eviction benchmarks, a per-head budget precomputed offline on pretraining text, combined with a training-free token score, shows competitive performance with methods that recompute the budget on every document or train it per target. Code available at https://github.com/Flegyas/marginal-attention
Problem

Research questions and friction points this paper is trying to address.

language models
attention mechanism
representation similarity
marginal attention
KV cache eviction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Marginal Attention Space
Representation Similarity
Input-Output Jacobian
KV Cache Eviction
Emergent Structure
🔎 Similar Papers
No similar papers found.