Score
Design and implement models, training objectives, and fine‑tuning procedures that predict, align, or transfer embeddings and supervisory signals across different data modalities, including multi‑task cross‑modal objectives, embedding prediction, and joint boundary-and-alignment training to produce unified cross‑modal representations. Analyze and evaluate the resulting representations and transfer behavior, ensuring cross‑modal supervision improves or at minimum does not degrade performance relative to unimodal baselines.
This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.
This work addresses the unclear interaction between feature alignment and target fitting in cross-modal fine-tuning, which often leads to a mismatch between feature-label structures across source and target domains, thereby degrading generalization. For the first time, this study theoretically characterizes their relationship by introducing the notion of “feature-label distortion,” and establishes a provable generalization bound on target error. Based on this analysis, a principle for joint optimization of alignment and fitting is derived. The resulting framework offers interpretable and actionable design guidelines for cross-modal fine-tuning. Extensive experiments demonstrate that the proposed method significantly outperforms current state-of-the-art approaches across multiple benchmark datasets, confirming its effectiveness and broad applicability.
This work addresses the lack of a systematic understanding of when cross-modal alignment (CA) and cross-modal prediction (CP) are effective in multimodal learning—a gap that often leads to suboptimal performance or even degradation relative to unimodal baselines. The authors propose a unified linear analytical framework under a structured signal–noise model with correlated interference, revealing complementary failure mechanisms of CA and CP. They introduce the first multimodal “phase diagram,” which delineates four distinct regimes: both methods succeed, only alignment works, only prediction works, or neither is effective. Leveraging separation ratio analysis, a unidirectional whitening mechanism, and a few-shot label-guided localization algorithm, this phase diagram enables practical guidance for method selection on real-world data. Experiments across synthetic, stereo vision, image–text, and astrophysical datasets validate its efficacy in identifying harmful multimodal configurations, offering a diagnostic tool for practitioners prior to model deployment.
This work addresses the challenge of efficiently equipping vision-language models with continuously evolving domain-specific skills, a task hindered by the high cost of conventional fine-tuning. The authors propose injecting capabilities from domain-specialized large language models into vision-language models through model fusion, enabling cross-modal skill transfer without additional training data or substantial computational resources. The study presents the first systematic analysis of the method’s applicability, fusion strategies, and hyperparameter sensitivity, offering quantitative evaluations of techniques such as Task Arithmetic (TA) and DARE across heterogeneous architectures. Experimental results demonstrate strong performance in instruction-following and cross-lingual tasks, while revealing limitations in mathematical reasoning, thereby delineating the effective boundaries and critical tuning factors for cross-modal skill injection.
This work investigates how vision-language models (VLMs) construct modality-agnostic task representations. We introduce the concept of *task vectors*—one-dimensional linear embeddings that compactly encode semantically equivalent task specifications across modalities (images/text) and formats (instructions/examples), outperforming full prompt inputs in downstream performance. Methodologically, we leverage autoregressive VLMs and employ cross-modal transfer evaluation, task vector disentanglement, and linear probing analysis. Key contributions: (1) We provide the first empirical evidence that task vectors exhibit strong generalization—enabling zero-shot transfer from LLMs to VLMs; (2) they can be extracted solely from instruction-based prompts without demonstrations; (3) they support cross-modal task triggering (e.g., image-to-text or text-to-image); and (4) they significantly improve zero-shot generalization across diverse tasks and VLM architectures. Our findings suggest that task semantics in VLMs are linearly separable and highly portable, offering a unified representation for multimodal task specification.
Multimodal large models (e.g., VLMs) face a novel attack surface where adversarial visual inputs bypass text-based safety mechanisms. Method: We propose a paradigm shift—performing *text-only unlearning* to achieve cross-modal safety alignment, eliminating reliance on multimodal retraining. Our approach introduces a language-space fusion architecture and a text-side parameter unlearning algorithm that operates exclusively in the textual modality, without requiring any visual training data. Contribution/Results: Empirical evaluation across six benchmark datasets shows our method reduces attack success rates to 2%–8%, while preserving original model functionality. Compared to multimodal fine-tuning baselines, it reduces computational overhead by 6×. This work is the first to demonstrate that pure text-space unlearning can effectively generalize to cross-modal safety alignment—challenging and extending prevailing safety training paradigms.
This study addresses the challenge of enabling models to learn efficiently and achieve strong performance in a specific test environment without relying on large-scale external data. The authors propose Test-Space Training (TST), a novel approach that systematically explores the feasibility of cross-modal self-supervised pretraining using only multimodal sensor data collected from devices within the target test environment. By leveraging multimodal alignment and environment-specific modeling, TST substantially reduces dependence on massive internet-scale datasets and demonstrates that modality diversity can partially compensate for limited data scale. Experimental results show that models trained exclusively on in-environment data attain performance on multiple downstream tasks comparable to that of prominent models such as DINOv2 and CLIP, which are pretrained on vast generic datasets.
This work addresses the unclear transferability mechanism between image understanding and generation tasks in existing unified multimodal models. It presents the first systematic investigation into cross-task transfer patterns between these two objectives and reveals that a shared Transformer backbone combined with a unified visual encoder enables stable knowledge transfer. Building on this insight, the study proposes a novel paradigm that enhances generative capabilities indirectly through training on understanding tasks, effectively mitigating distribution shift issues. The approach is validated on three critical competencies—counting, spatial reasoning, and text recognition/generation—demonstrating significant improvements in generation performance without compromising visual fidelity.
This work systematically investigates the core mechanisms of modality interaction in multimodal pretraining, focusing on knowledge transfer, synergistic effects, and fusion timing. Through experiments on both synthetic and large-scale real-world datasets, it provides the first empirical evidence of asymmetric cross-modal knowledge flow and demonstrates that data complexity governs whether modalities exhibit synergy or competition. The study further validates that early unified fusion consistently outperforms late alignment. Leveraging an architecture featuring shared attention and normalization layers with modality-specific feedforward components, the proposed approach is evaluated on a 13.5B mixture-of-experts model trained on 2 trillion tokens, confirming its effectiveness. Additionally, the paper introduces a highly efficient pretraining strategy that achieves strong generative performance using only 5% of the typical computational budget.
This work investigates whether the understanding and generation branches of unified multimodal models share a transferable semantic space. To this end, the authors propose a cross-branch semantic guidance framework that extracts intervention-based semantic directions from the understanding branch and transfers them to the generation branch for controllable image synthesis. Their experiments reveal, for the first time, a semantic asymmetry between the two branches: the understanding branch encodes object-level semantics, whereas the generation branch relies more heavily on low-level appearance features. The proposed method enables effective semantic transfer from understanding to generation, significantly improving the semantic fidelity of generated images; however, reverse transfer yields limited gains, demonstrating that architectural unification does not inherently guarantee semantic alignment.
This study investigates the efficacy of unimodal forgetting transfer to the other modality in vision-language models and its robustness under typographic attacks. It reveals, for the first time, that cross-modal forgetting exhibits asymmetry and fragility. To address this, the authors propose CrossInf, an influence function–based method that selectively targets critical Transformer modules to enhance forgetting transfer. Through comprehensive validation—including influence function analysis, CKA representation alignment, human evaluation (κ=0.77), and simulated typographic attacks—CrossInf reduces the forgetting transfer gap by over 50% in strongly fused architectures while driving typographic attack success rates nearly to zero, thereby substantially improving both model security and utility preservation.