Score
Designs and implements algorithms, model architectures, and training methods that condition model behavior on signals from multiple modalities—including contextual, cross-modal, task-specific, probabilistic, and temporal cues—and that fuse those signals into coherent internal representations. This competence covers specifying conditioning mechanisms (e.g., attention, gating, latent-variable conditioning), temporal alignment and fusion strategies, and analyzing how different probabilistic or task-conditioned fusion choices affect downstream performance.
This paper addresses the challenges in developing general-purpose multimodal large language models (MLLMs) capable of cross-modal coordination across six generative modalities: text, image, music, video, human motion, and 3D objects. To this end, it proposes a novel unified architecture integrating Transformer-based and diffusion-based paradigms, augmented with self-supervised learning (SSL), mixture-of-experts (MoE), reinforcement learning from human feedback (RLHF), and chain-of-thought (CoT) reasoning. The work introduces the first taxonomy covering all six modalities and identifies shared enabling mechanisms for cross-modal transfer. It further argues that structured reasoning and modular decoupling are critical to improving interpretability and generalization. The resulting comprehensive MLLM technology landscape clarifies common bottlenecks and transferable methodologies, providing both theoretical foundations and practical guidelines for building universal, adaptive, and interpretable multimodal systems.
Current multimodal fusion methods in AI predominantly employ static weighting strategies, overlooking the neuroscientific principle of inverse effectiveness—the phenomenon wherein weaker unimodal cues elicit stronger multimodal integration gains. To address this, we propose Inverse-Effectiveness-Driven Multimodal Fusion (IEMF), the first approach to formalize this principle as a learnable, differentiable dynamic fusion mechanism that adaptively modulates integration strength based on unimodal confidence estimates. The IEMF module integrates attention mechanisms, dynamic gating, and cross-modal confidence estimation, and is compatible with both artificial neural networks (ANNs) and spiking neural networks (SNNs). Evaluated on audio-visual classification, continual learning, and multimodal question answering, IEMF significantly enhances model robustness—particularly under noisy or sparse inputs—while reducing computational overhead by up to 50%. Moreover, it demonstrates strong generalization across diverse tasks and modalities.
This study investigates cross-modal decision-making mechanisms in vision-language models (VLMs) under modality conflicts—e.g., an image of a dog paired with the caption “This is a cat.” We systematically construct conflict-rich multimodal samples and employ attention head localization, representation space analysis, and instruction-guided modality selection tasks. Our analysis reveals an intrinsic modality bias in VLMs and identifies two functionally distinct architectural components: (i) dedicated attention heads that regulate modality preference, and (ii) transferable, modality-agnostic “router heads” that dynamically route information across modalities. Crucially, targeted intervention on router heads significantly improves model accuracy in detecting multimodal consistency. This work provides the first empirical evidence of a hierarchical modality fusion architecture within VLMs, uncovering interpretable, controllable mechanisms for multimodal reasoning and cross-modal alignment.
The computational principles underlying Transformer attention remain poorly understood. Method: We establish, for the first time, a theoretical correspondence between attention and Pavlovian conditioning—mapping queries, keys, and values to test stimuli, conditioned stimuli, and unconditioned stimuli, respectively—and formalize attention as Hebbian-driven formation of transient associative memory. Leveraging a linear attention model, we integrate associative learning theory, matrix analysis, and error propagation analysis to rigorously characterize this memory process. Contribution/Results: We derive a fundamental single-head capacity bound of *O*(√*dₖ*), where *dₖ* is the key dimension, and uncover inherent trade-offs among model depth, width, and attention head redundancy. Furthermore, we propose a biologically plausible learning rule grounded in this framework, offering a novel theoretical foundation for designing efficient and neuroscientifically credible Transformer architectures.
This work investigates the syntactic modeling capabilities of multimodal language models (MLLMs), specifically examining how visual context induces syntactic priming effects. To this end, we introduce PRISMATIC—the first large-scale multimodal structural priming dataset—and propose a reference-free metric for quantifying syntactic priming strength. Through systematic comparison of dual-encoder versus fusion-encoder architectures, we make the first observation that robust positive correlation between visual similarity and syntactic priming strength emerges exclusively in fusion-encoder models—a pattern strongly aligned with cross-modal coupling mechanisms documented in human psycholinguistics; no such correlation is found in dual-encoder models. These findings indicate that fusion architectures achieve deeper syntactic–visual representational coupling, offering novel empirical evidence for structured cognitive modeling in MLLMs and establishing an interpretable, task-grounded evaluation paradigm for syntactic priming in multimodal settings.
Existing surveys predominantly examine isolated components of multimodal pipelines and lack empirically grounded, pedagogically oriented integration frameworks for teaching and learning contexts. Method: This study introduces the first taxonomy and analytical framework covering five core modalities—natural language, video, sensor data, human-centered signals, and environmental logs—and proposes a novel “mid-fusion” paradigm for multimodal data integration. It further innovates by applying citation graph pruning to achieve structured, high-precision literature synthesis. Contribution/Results: Through systematic review, taxonomic modeling, and multimodal fusion design, we demonstrate that multimodal synergy enables detection of fine-grained learning behaviors imperceptible to unimodal analysis. While prediction accuracy remains largely unchanged, interpretability improves significantly, yielding deeper insights into learners’ cognitive-affective states and training outcomes.
This work systematically investigates the core mechanisms of modality interaction in multimodal pretraining, focusing on knowledge transfer, synergistic effects, and fusion timing. Through experiments on both synthetic and large-scale real-world datasets, it provides the first empirical evidence of asymmetric cross-modal knowledge flow and demonstrates that data complexity governs whether modalities exhibit synergy or competition. The study further validates that early unified fusion consistently outperforms late alignment. Leveraging an architecture featuring shared attention and normalization layers with modality-specific feedforward components, the proposed approach is evaluated on a 13.5B mixture-of-experts model trained on 2 trillion tokens, confirming its effectiveness. Additionally, the paper introduces a highly efficient pretraining strategy that achieves strong generative performance using only 5% of the typical computational budget.
This work addresses the challenge of efficiently equipping vision-language models with continuously evolving domain-specific skills, a task hindered by the high cost of conventional fine-tuning. The authors propose injecting capabilities from domain-specialized large language models into vision-language models through model fusion, enabling cross-modal skill transfer without additional training data or substantial computational resources. The study presents the first systematic analysis of the method’s applicability, fusion strategies, and hyperparameter sensitivity, offering quantitative evaluations of techniques such as Task Arithmetic (TA) and DARE across heterogeneous architectures. Experimental results demonstrate strong performance in instruction-following and cross-lingual tasks, while revealing limitations in mathematical reasoning, thereby delineating the effective boundaries and critical tuning factors for cross-modal skill injection.
Existing approaches struggle to finely dissect the interaction mechanisms among modalities in multimodal large language models and their task-specific dependencies. This work introduces Partial Information Decomposition (PID) into the analysis of multimodal foundation models, proposing the Sensory PID framework to quantify, at the decision level, the unique, redundant, and synergistic contributions of visual, auditory, and linguistic inputs—extending PID to tri-modal systems for the first time. The study reveals general patterns of modality usage: reasoning and localization tasks heavily rely on cross-modal synergy, whereas knowledge-intensive tasks predominantly depend on language. It also uncovers a vision-dominated information bottleneck in audio-visual fusion. Building on these insights, the authors devise a PID-guided weight reweighting strategy that yields preliminary performance gains in multimodal reasoning and localization.
This study investigates how vision-language models integrate visual and textual information during chain-of-thought (CoT) reasoning and examines their susceptibility to misleading textual cues. Through dynamic tracking of CoT confidence trajectories, controlled interventions introducing deceptive text, and comparative analysis across 18 models, the work reveals a previously undocumented “answer inertia” phenomenon: models exhibit strong reliance on textual cues even when capable of correction. The findings indicate that while reasoning-focused training enhances corrective capacity, it fails to eliminate textual bias. Although instruction-tuned models less frequently cite misleading information explicitly, their reasoning traces more readily expose inconsistencies between visual and textual inputs. These results suggest that although CoT partially reflects multimodal integration, its apparent fluency may mask an implicit overreliance on textual signals.
This study addresses the challenge that multimodal models struggle to reliably bind textual prompts (e.g., “image”) to their corresponding input modalities, resulting in inadequate source modality tracking. The work formally defines and empirically investigates the “source modality monitoring” problem for the first time, introducing an evaluation paradigm grounded in target-modality information retrieval. By integrating syntactic manipulation with semantic perturbation, the authors systematically assess binding mechanisms across eleven prominent vision-language models. Their findings reveal that semantic cues dominate the binding process when modality distributions exhibit significant divergence, consistently outweighing syntactic signals. These insights offer critical evidence and a novel perspective for enhancing the reliability and robustness of multimodal agents.