Strong Helps Weak: Directional Cross-Modal Alignment Transfer in Multi-modal LLMs

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of high enhancement costs for data-scarce modalities and the difficulty of acquiring intra-modal variants for model merging in multimodal large language models. To this end, we propose DCAT, a framework that leverages strongly aligned source modalities to enhance the text-alignment capability of weak target modalities. Our method reveals an asymmetric gain phenomenon and derives a theoretical lower bound on mutual information, enabling closed-form solutions in weight space and cross-modal alignment transfer via a small calibration set. Experimental results demonstrate that DCAT significantly outperforms existing model merging approaches, effectively improving multimodal understanding performance for data-scarce modalities without requiring additional fine-tuning.
📝 Abstract
Multi-modal large language models (MLLMs) achieve strong modality understanding by pairing a large language model (LLM) with an encoder for a target modality such as vision, video, or audio. However, improving an MLLM's capability for a given modality typically requires additional training on large modality-specific datasets, incurring substantial data collection and compute costs. Model merging offers an alternative, but it is often infeasible for data-scarce, large per-sample size, or domain-specific modalities (\textit{e.g.}, audio and video), where same-modality model variants are rarely available. In this work, we characterize an intriguing asymmetric phenomenon: merging a well-aligned, data-rich source-modality MLLM into a data-scarce target-modality MLLM substantially improves the target on its own benchmarks. Our theoretical and empirical analyses show that this gain stems from enhanced alignment between modality-specific and textual tokens, induced by the stronger donor modality. Specifically, we derive a mutual-information lower bound that is monotonic in alignment-related quantities and strongly correlated with downstream MLLM performance. Building on this principle, we propose Directional Cross-modal Alignment Transfer (DCAT), a novel framework that transfers textual alignment from a strong, well-aligned source (donor) modality to a weak target (recipient) modality, boosting target-modality performance without further fine-tuning. We further show that the alignment-enhancing objective admits a closed-form weight-space solution computed from only a small calibration set. DCAT outperforms existing model-merging methods, offering an efficient path toward cross-modal alignment transfer. Project page with code is available at \url{https://seohoiki3215.github.io/DCAT_project_page}
Problem

Research questions and friction points this paper is trying to address.

Multi-modal LLMs
Cross-Modal Alignment
Data-scarce Modalities
Model Merging
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Modal Alignment Transfer
Multi-modal Large Language Models
Model Merging
Mutual Information Lower Bound
Directional Cross-modal Alignment Transfer (DCAT)
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hoigi Seo
Dept. of ECE, Seoul National University, Republic of Korea
B
Byung Hyun Lee
Dept. of ECE, Seoul National University, Republic of Korea
Minjun Kim
Minjun Kim
Ph.D. student at Seoul National University
Graph MiningModel Compression
D
Dohyun Mah
Dept. of ECE, Seoul National University, Republic of Korea
Jongho Lee
Jongho Lee
TEAMREBOOTT Inc.
large language modelsnatural language processing
Se Young Chun
Se Young Chun
Department of Electrical and Computer Engineering, Seoul National University
computational imagingmachine learningsignal processingmultimodal processing