Score
Designs and applies methods to estimate and analyze the entropy of attention distributions produced by attention mechanisms (e.g., transformer attention heads). This includes building estimators for attention-head entropy, probing individual heads to identify condition-dependent differences, performing head ablations or interventions to test causal roles, and comparing entropy measures across experiments to detect transfer or generalization effects.
This work addresses the substantial bias and poor generalizability of transfer entropy (TE) estimation under stationary processes. We propose TREET, a Transformer-based neural estimator that innovatively integrates the Donsker–Varadhan dual representation with self-attention mechanisms, establishing the first differentiable TE optimization framework grounded in the functional representation lemma. This unified framework supports TE estimation, channel capacity computation, and density inversion. By circumventing the high sensitivity of conventional nonparametric estimators to high-dimensional, small-sample settings, TREET achieves significant performance gains over state-of-the-art TE estimators on standard benchmarks. It is the first method to successfully estimate channel capacity for memory channels and demonstrates robust causal inference and joint density modeling on real-world physiological data from sleep apnea patients.
The internal information distribution mechanism of Transformer models remains poorly understood. Method: This paper proposes an information-entropy-based probing method to quantify token-level uncertainty and model the dynamic entropy evolution across layers, without requiring additional training or labeled supervision. Contribution/Results: Applying systematic entropy trajectory analysis to GPT-family models, we uncover, for the first time, an alternating pattern of information compression and diffusion between feed-forward layers and attention heads, identifying critical information bottlenecks and representation transition pathways. The method offers strong interpretability and establishes the first lightweight, information-flow-oriented analytical framework for Transformer interpretability research. Furthermore, it enables the design of novel model evaluation metrics grounded in entropy dynamics—providing a principled, quantifiable basis for assessing representational efficiency and layer-wise information processing in large language models.
This work addresses the limited trustworthiness of Transformer models in high-stakes applications, which stems from insufficient understanding of their internal decision-making mechanisms. To bridge this gap, we propose a mechanistic interpretability approach based on targeted interventions on attention heads, integrating causal analysis with neural circuit probing to systematically uncover the model’s decision processes and underlying cognitive mechanisms. Our method substantially enhances the interpretability of Transformer internals and offers an innovative pathway toward the design and control of highly reliable AI systems, while also enabling the discovery of novel scientific insights encoded within these models.
The opaque internal computation mechanisms of Transformers hinder their interpretability and robustness evaluation. To address this, we propose the first Shannon entropy-based framework for dynamic residual flow analysis, modeling inter-layer information entropy evolution as a universal “computational signature.” Our method requires no architectural priors and is applicable across multimodal Transformer architectures—including LLMs and ViTs—under black-box conditions. It achieves three key functionalities: (1) precise discrimination among mainstream model families; (2) automatic classification of task prompt types; and (3) accurate prediction of output accuracy—demonstrating strong correlation (ρ > 0.82) between entropy trajectories and accuracy across multiple benchmarks. Crucially, this work establishes information entropy evolution as a generalizable and transferable computational representation paradigm for Transformers—a novel contribution to foundational model analysis.
This work addresses the poor conditioning of the Jacobian matrix in Transformer attention mechanisms, which often leads to training instability and performance degradation. For the first time, it explicitly establishes a theoretical link between the condition number of the attention Jacobian and the spectral properties of the query, key, and value projection matrices. Building on this insight, the paper proposes a general, plug-and-play spectral regularization strategy that improves the Jacobian’s condition number by optimizing the singular value distribution of these projection matrices. Notably, the method requires no architectural modifications and consistently enhances performance across diverse Transformer variants and tasks, demonstrating both its effectiveness and broad applicability.
This study addresses the propensity of small language models (1B–1.7B parameters) to generate high-confidence errors and hallucinations under resource-constrained conditions, where the relationship between their internal dynamics and factual accuracy remains unclear. Leveraging the TruthfulQA benchmark, the work conducts a token-level analysis of entropy, attention patterns, and hidden state evolution, proposing for the first time a dynamic classification of models into deterministic, exploratory, and balanced types based on entropy trajectories. It reveals structured associations among these types, output truthfulness, attention distributions, and representational pathways. The findings demonstrate that high factual accuracy arises from orderly entropy and attention dynamics, establishing an interpretable and optimizable paradigm of internal uncertainty for designing low-hallucination, high-reliability edge-deployed small models.
This study addresses the fragility of causal inference in attention head ablation, which often stems from semantic bias in interventions, metric saturation, and the absence of controls. Focusing on GPT-2, this work compares pre- and post-projection ablation differences to reveal projection-level confounding effects. It introduces a continuous log-probability metric to mitigate saturation and constructs matched random heads as control baselines, with evaluations conducted via Spearman correlation and Monte Carlo testing. This research establishes the necessity of non-saturating metrics and matched controls for robust causal inference. The corrected head importance rankings demonstrate high stability across data splits (ρ=0.974), with top-5 heads significantly outperforming the control distribution; however, evidence for task specificity remains inconclusive.
Existing interpretability methods often erroneously attribute specific computational roles to attention heads without verifying their generalization across diverse prompts. This work proposes the KID role classification framework and a three-stage analysis pipeline that combines activation patching with same-answer control conditions to expose widespread pseudo-semantic specificity in conventional attribution approaches. By integrating capability-selective screening (CSS), singular value decomposition (SVD), and activation transduction under matched controls, we systematically evaluate attention head functionality across multiple 7–8B instruction-tuned models. Our findings demonstrate that the majority of attention heads previously identified by standard methods as performing specific roles fail to consistently transfer their purported computational functions across different prompts, thereby challenging the dominant attribution paradigm in mechanistic interpretability.
Existing approaches to long-context large language model inference often rely on fixed sparsity patterns or uniform computational budgets, overlooking the dynamic disparities among attention heads and across context positions. This work proposes EntropyInfer, a training-free framework that adaptively partitions attention heads into rigid and dynamic categories during the prefill phase based on attention entropy, enabling context-aware allocation of computational resources. During decoding, it introduces an untrained KV cache compression mechanism that preserves critical cached information aligned with the generated content. EntropyInfer achieves fine-grained, context-adaptive inference acceleration without requiring model retraining. Evaluated on Llama, Qwen, and openPangu models, it delivers up to 2.39× end-to-end speedup on sequences exceeding 100k tokens while incurring minimal quality degradation, substantially outperforming baselines such as SnapKV and AdaKV.
This study addresses the challenge of identifying attention-head circuits responsible for specific tasks within pretrained Transformers under fully unsupervised conditions. The authors propose a three-step methodology: first, attention heads are ranked without supervision using a time-integrated participation ratio derived from spectral signals; second, candidate circuits are selected by aligning with task-specific activation patterns; and third, causal relevance is validated through grouped ablation against randomized controls. This approach is the first to identify causal circuits without requiring labels or gradients, consistently uncovering 2–6 essential inductive circuits across models ranging from 51M to 7B parameters and diverse architectures—ablating which degrades performance by 94%–100%. The work further reveals that the proportion of specialized computational heads remains conserved across models, falling within a narrow range of 17%–19%.