Rethinking Cross-Modal Fine-Tuning: Optimizing the Interaction between Feature Alignment and Target Fitting

πŸ“… 2026-01-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the unclear interaction between feature alignment and target fitting in cross-modal fine-tuning, which often leads to a mismatch between feature-label structures across source and target domains, thereby degrading generalization. For the first time, this study theoretically characterizes their relationship by introducing the notion of β€œfeature-label distortion,” and establishes a provable generalization bound on target error. Based on this analysis, a principle for joint optimization of alignment and fitting is derived. The resulting framework offers interpretable and actionable design guidelines for cross-modal fine-tuning. Extensive experiments demonstrate that the proposed method significantly outperforms current state-of-the-art approaches across multiple benchmark datasets, confirming its effectiveness and broad applicability.

Technology Category

Machine Learning: Multimodal LearningComputer Vision: Multi-modal VisionNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: User modeling for targeted and personalized online advertisingWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web data
πŸ“ Abstract
Adapting pre-trained models to unseen feature modalities has become increasingly important due to the growing need for cross-disciplinary knowledge integration. A key challenge here is how to align the representation of new modalities with the most relevant parts of the pre-trained model's representation space to enable accurate knowledge transfer. This requires combining feature alignment with target fine-tuning, but uncalibrated combinations can exacerbate misalignment between the source and target feature-label structures and reduce target generalization. Existing work, however, lacks a theoretical understanding of this critical interaction between feature alignment and target fitting. To bridge this gap, we develop a principled framework that establishes a provable generalization bound on the target error, which explains the interaction between feature alignment and target fitting through a novel concept of feature-label distortion. This bound offers actionable insights into how this interaction should be optimized for practical algorithm design. The resulting approach achieves significantly improved performance over state-of-the-art methods across a wide range of benchmark datasets.
Problem

Research questions and friction points this paper is trying to address.

cross-modal fine-tuning
feature alignment
target fitting
generalization
feature-label distortion
Innovation

Methods, ideas, or system contributions that make the work stand out.

feature-label distortion
generalization bound
cross-modal fine-tuning
feature alignment
target fitting
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
T
Trong Khiem Tran
Washington State University
M
M. Dao
National University of Singapore
P
Phi Le Nguyen
Hanoi University of Science and Technology
T
T. Truong
National Institute of Advanced Industrial Science and Technology
T
T. Hoang
Washington State University