Score
Designs and implements attention modules that compute and apply weighted interactions between heterogeneous feature streams, aligning tokens across modalities and across time or space using cross-attention, asymmetric attention, or mixture-of-transformers routing. Includes context- or prompt-conditioned and guided cross-attention variants that modulate information flow to disambiguate signals, emphasize task-relevant steps, capture long-range inter-modal dependencies, and support explicit correspondence prediction and evaluation.
This paper addresses the challenges in developing general-purpose multimodal large language models (MLLMs) capable of cross-modal coordination across six generative modalities: text, image, music, video, human motion, and 3D objects. To this end, it proposes a novel unified architecture integrating Transformer-based and diffusion-based paradigms, augmented with self-supervised learning (SSL), mixture-of-experts (MoE), reinforcement learning from human feedback (RLHF), and chain-of-thought (CoT) reasoning. The work introduces the first taxonomy covering all six modalities and identifies shared enabling mechanisms for cross-modal transfer. It further argues that structured reasoning and modular decoupling are critical to improving interpretability and generalization. The resulting comprehensive MLLM technology landscape clarifies common bottlenecks and transferable methodologies, providing both theoretical foundations and practical guidelines for building universal, adaptive, and interpretable multimodal systems.
Contemporary multimodal large language models (MLLMs) suffer from inconsistent cross-modal attention and progressive layer-wise attenuation, hindering fine-grained perception, cognition, and affective understanding in advanced multimodal tasks. To address these limitations, we propose Modular Dual-path Attention (MODA), a novel architecture featuring three key innovations: (1) a “correct-post-alignment” strategy that decouples modality alignment from cross-layer token mixing; (2) adaptive masked attention enabling modality-specific interaction patterns; and (3) a unified bimodal embedding space constructed via foundational vector mapping. MODA preserves semantic fidelity while enhancing cross-modal coherence across layers. We comprehensively evaluate MODA on 21 diverse multimodal benchmarks—including visual reasoning, emotion recognition, and compositional understanding—demonstrating consistent and significant improvements over state-of-the-art MLLMs. All code and interactive demos are publicly released.
This work addresses the challenge of enhancing neural networks’ ability to focus on salient information in long-sequence and multimodal tasks. By establishing a unified theoretical framework for attention mechanisms, the study systematically analyzes their mathematical foundations, computational properties, and cross-task generalizability. The framework is instantiated across diverse architectures—including autoregressive Transformers, bidirectional encoders, Vision Transformers, and cross-modal attention models—demonstrating consistent performance gains. The research further uncovers an intrinsic relationship between attention structure and model interpretability, validates empirical scaling laws governing training dynamics and performance, and achieves state-of-the-art results on multiple benchmark datasets. Attention visualization techniques are employed to enhance model transparency, offering insights into the decision-making process of these architectures.
This paper addresses the poor generalizability, training instability, and weak interpretability of conventional spatial attention mechanisms in convolutional neural networks (CNNs), which stem from irregular, pixel-level attention regions. To this end, we propose a parametric Rectangular Spatial Attention Module (RSAM) that explicitly defines a rectangular attention region using only five learnable parameters. RSAM is fully differentiable and enables end-to-end joint optimization, serving as a plug-and-play component compatible with arbitrary CNN architectures. Our key contribution is the first explicit geometric constraint of spatial attention to a rectangle—enhancing boundary regularity, training stability, and cross-sample generalization, while improving semantic interpretability of attended locations. Extensive experiments on multiple benchmarks demonstrate that RSAM consistently outperforms pixel-wise attention methods, achieving significant gains in classification accuracy, robustness to input perturbations, and visual localization consistency.
Existing Transformer interpretability research predominantly focuses on MLP neurons and simple factual concepts, neglecting attention mechanisms and lacking a unified analytical framework for complex, abstract concepts. Method: We propose Concept-Agnostic Attention Module Discovery and Scalar Intervention (SAMD/SAMI), the first method to directly map arbitrary complex concepts—e.g., ethical judgments or reasoning steps—to specific attention heads and modulate their influence via a single scalar parameter. Our approach leverages concept vectorization and cosine similarity ranking, enabling consistent cross-modal and cross-task analysis. Contribution/Results: SAMD/SAMI demonstrates robust module localization stability before and after LLM post-training; achieves a 72.7% reduction in jailbreaking success rate on HarmBench and a 1.6% absolute accuracy gain on GSM8K; and generalizes successfully to Vision Transformers, where it effectively suppresses ImageNet classification accuracy—validating both its broad applicability and precise controllability.
This work addresses the lack of a systematic understanding of how attention mechanisms in Vision Transformers jointly process positional and content information. The authors propose a Bilinear Factorization Decomposition (BFD) framework that, for the first time, achieves statistical disentanglement of positional and content factors through ANOVA decomposition, combined with singular value decomposition (SVD) of the QK^T matrix to uncover dominant interaction modes within attention. Their analysis reveals that attention energy is primarily driven by content-content interactions; DINOv2 exhibits stronger content-position coupling and a richer distribution of interaction modes; and intermediate layers enhance shape perception by jointly preserving positional structure and amplifying semantic signals.
This work addresses the limitations of conventional multi-head attention in permutation-invariant set tasks, where fixed value projections hinder the modeling of complex input-dependent target mappings. To overcome this, the authors propose a context-adaptive attention mechanism that dynamically adjusts value projections using a family of matrix zonotopes—defined as a center matrix plus a weighted sum of generator matrices gated by the input—thereby enhancing representational capacity while preserving permutation equivariance. They introduce a Transform Degrees of Freedom (TDOF) metric and theoretically demonstrate that a single layer of the proposed mechanism can efficiently represent targets with high TDOF, circumventing the need for deep stacking inherent in traditional approaches. Empirical results show significant performance gains over standard attention on high-rank sparse combinatorial set prediction tasks, while matching its performance on aggregate statistical tasks.
This work addresses the limited understanding of the dynamic mechanisms governing when multimodal large language models (MLLMs) invoke visual versus textual information during generation. It presents the first systematic analysis of token-by-token attention dynamics, employing attention trajectory tracing, causal masking interventions, and test-time modulation across multiple open-source models. The study reveals consistent patterns: visual attention peaks precisely at tokens semantically related to the image, while instruction tokens are revisited during task transitions. Building on these insights, the authors propose an attention-guided intervention strategy that significantly enhances performance on multimodal tasks. Furthermore, controlled perturbations expose a critical vulnerability—models readily degrade to relying solely on linguistic priors or generate cross-modal inconsistencies—thereby underscoring the pivotal role of attention mechanisms in coherent multimodal generation.
Standard dot-product attention incurs computational and memory bottlenecks in long-context scenarios due to dense pairwise interactions. This work proposes Gaussian Mixture Attention (GMA), which maps queries and keys into a shared latent routing space and implicitly computes similarities through K learnable Gaussian mixture components, while reading from and writing to a K-slot latent memory—thereby avoiding explicit construction of the N×N attention matrix. GMA achieves linear-time sequence mixing via probabilistic latent-variable routing, offering interpretable responsibility assignment, non-negative low-rank approximation, and stable local routing. Its end-to-end differentiable design supports both causal and bidirectional variants, enabling linear memory scaling with fixed K. Empirically, GMA matches standard attention in long-context classification, outperforms several linear methods on WikiText-103 in its causal form, and exhibits broad, semantically aligned component utilization as revealed by responsibility analysis.