Score
Designs and implements transformer-based architectures that integrate and fuse feature representations across multiple spatial and semantic scales and across modalities, producing unified transformer embeddings or feature maps for downstream tasks. Involves selecting attention and connectivity patterns for hierarchical multimodal fusion, engineering integration modules that produce task-ready representations (e.g., segmentation-ready features) while controlling computational and memory cost.
Existing research lacks a systematic analysis of modality–language backbone integration mechanisms in multimodal large language models (MLLMs). Method: We systematically survey 125 MLLMs published between 2021 and 2025, proposing the first large-language-model-centric three-dimensional taxonomy—architectural integration, representation learning, and training paradigms—to unify cross-modal alignment pathways. Leveraging bibliometric analysis, architectural decoupling, and fine-grained modeling of embeddings and loss functions, we characterize evolutionary patterns in fusion granularity, joint representation design, and objective function development. Contribution/Results: This work fills a critical theoretical gap in structured analysis of MLLM fusion mechanisms, identifies shared bottlenecks—including misaligned modality-specific feature hierarchies and suboptimal cross-modal optimization objectives—and delivers a reusable theoretical framework and practical guidelines for developing robust, scalable multimodal foundation models.
Vision Transformers (ViTs) suffer from weak hierarchical modeling capability in conventional Patch Merging, struggling to jointly capture global dependencies and local details. To address this, we propose a brain-inspired Stepwise Patch Merging (SPM) paradigm. SPM decouples global integration from local refinement: a learnable Multi-Scale Aggregation (MSA) module formalizes the brain’s multi-scale fusion mechanism to enhance long-range modeling; a Guided Local Enhancement (GLE) module dynamically strengthens salient local structures. Integrated into hierarchical ViT backbones, SPM supports end-to-end joint training. Extensive experiments on ImageNet-1K, COCO, and ADE20K demonstrate that SPM significantly improves performance on dense prediction tasks—achieving state-of-the-art accuracy and robustness in object detection and semantic segmentation.
Transformer positional encodings and attention mechanisms have long lacked a unified geometric and physical interpretation. Method: This paper introduces the first framework embedding Transformers within geometric field theory: discrete token positions are mapped to a continuous embedding manifold, and self-attention is formalized as a kernel-modulated integral operator defined on this manifold. By integrating manifold embedding, differential geometry, and field-theoretic principles, attention is recast as function modulation and transformation in continuous space. Contribution/Results: The framework provides an interpretable geometric semantics for core Transformer components—unifying the mathematical foundations of positional encoding and attention—and establishes a theoretical bridge between discrete neural architectures and continuous field theory. It enables principled design of next-generation attention mechanisms endowed with explicit geometric priors, advancing both interpretability and inductive bias engineering in deep learning.
This work proposes a safety-compliant multimodal Transformer architecture designed to meet automotive functional safety standards, addressing the lack of fault-tolerant and robust designs in existing Transformer models. By employing independent encoders to map heterogeneous sensor inputs into a shared latent space, the architecture structurally embeds redundancy and diversity at the representation level, enabling continuous operation under modality-level failures. This approach represents the first integration of multimodal foundation models with established automotive functional safety practices, facilitating certifiable autonomous driving systems. The model maintains consistent scene understanding even when certain modalities degrade, thereby offering a viable pathway for deploying Transformers in safety-critical applications.
Transformer models suffer from high computational overhead due to the Softmax-based self-attention mechanism, hindering their efficient deployment in vision tasks. To address this, we propose AttentioN-Free ViT (AF-ViT), the first Vision Transformer architecture that replaces self-attention entirely with gated linear units (GLUs) as the core building block—eliminating both the attention module and the second non-gated MLP layer in the standard ViT block. The resulting architecture is a purely MLP-based, attention-free vision model that significantly reduces FLOPs and memory footprint while preserving representational capacity. Extensive experiments demonstrate that AF-ViT achieves accuracy on par with standard ViTs across multiple benchmarks—including ImageNet classification, COCO object detection, and ADE20K semantic segmentation—while accelerating both training and inference by 35%–52%. These results validate the effectiveness and practicality of attention-free paradigms for visual representation learning.
This work investigates whether CNNs and Vision Transformers (ViTs) share unified learning mechanisms in image recognition, with a focus on the role of multi-head attention (MHA). To this end, we propose Single-Node Performance (SNP), a metric that quantifies the discriminative capability of individual feed-forward and MHA sub-module nodes toward label clusters, revealing common mechanisms of progressive signal enhancement and noise suppression. We further design the ANDC pruning method, achieving parameter-efficient compression without accuracy degradation. Notably, we discover— for the first time—spontaneous symmetry breaking among MHA heads, leading to head-wise label specialization, and establish a quantitative *modus vivendi* model characterizing their coexistence. Extensive experiments on CIFAR-100 and Flowers-102 validate the universality of these mechanisms, demonstrating both high accuracy and strong model compactness.
This work addresses the limited interpretability of heterogeneous attention mechanisms—such as co-attention—in multimodal or multi-source information fusion, where existing approaches struggle to elucidate their internal workings. To bridge this gap, the paper introduces the first general-purpose interpretability framework tailored specifically for heterogeneous attention. By integrating attention analysis, semantic interpretation, and logical reasoning, the proposed method establishes a unified analytical paradigm. The framework is successfully applied to representative Transformer-based models, enabling in-depth semantic and logical dissection of heterogeneous attention mechanisms. Experimental results demonstrate its broad applicability and practical utility across diverse architectures, offering new insights into how such attention modules process and integrate heterogeneous inputs.
This work addresses the limitations of existing approaches that predominantly rely on non-native late-fusion architectures and lack a systematic definition or unified framework for native multimodal modeling. The paper proposes a formal roadmap for Native Multimodal Modeling (NMM), introducing, for the first time, a rigorous formulation of “architectural nativeness.” Leveraging an input–output duality perspective, it categorizes models into three types: Multi-to-Text, Multi-to-Target, and Multi-to-Multi. A full-stack, industrial-grade NMM framework is developed, encompassing data governance, early/mid-stage fusion, end-to-end training, and deployment, all realized through a unified Transformer architecture that enables symbiotic cross-modal understanding and generation. This paradigm demonstrates superior performance and strong scalability across multiple tasks, offering a clear pathway toward truly native multimodal models.
This work addresses the lack of principled understanding in existing literature regarding the choice between cross-attention and feature concatenation strategies for multimodal fusion, which has largely relied on empirical heuristics. Through controlled experiments and theoretical analysis, we demonstrate for the first time that feature alignment quality is the key determinant of fusion strategy performance: under pre-aligned features, concatenation consistently outperforms cross-attention by 4.1–5.1 percentage points across all data scales, with its advantage becoming more pronounced as alignment degrades. Building on this insight, we develop a theoretical decision framework grounded in sample complexity and validate our findings using features extracted from ResNet-18 and CLIP ViT-B/32 on controlled datasets.
This work investigates the *fundamental universality* of Transformer architectures—specifically, their expressive capacity limits across diverse AI tasks and the minimal structural conditions required for universal approximation. Method: Integrating tools from function approximation theory, computational complexity analysis, and abstract architectural modeling, we establish the first systematic characterization of Transformers’ *structural minimality* and *approximation rates*, rigorously distinguishing theoretical prerequisites for robust generalization (e.g., continuous function approximation) versus fragile generalization (e.g., long-range dependency modeling). We further propose a *hierarchical universality framework* that quantifies how key design parameters—including number of attention heads, depth, and positional encoding schemes—govern expressive power. Contribution/Results: Our analysis provides foundational theoretical guarantees for provably reliable model compression, efficient lightweight architecture search, and generalization-aware design, bridging theoretical insights with practical deployment constraints in modern foundation models.