π€ AI Summary
This work addresses the unclear interaction between feature alignment and target fitting in cross-modal fine-tuning, which often leads to a mismatch between feature-label structures across source and target domains, thereby degrading generalization. For the first time, this study theoretically characterizes their relationship by introducing the notion of βfeature-label distortion,β and establishes a provable generalization bound on target error. Based on this analysis, a principle for joint optimization of alignment and fitting is derived. The resulting framework offers interpretable and actionable design guidelines for cross-modal fine-tuning. Extensive experiments demonstrate that the proposed method significantly outperforms current state-of-the-art approaches across multiple benchmark datasets, confirming its effectiveness and broad applicability.
π Abstract
Adapting pre-trained models to unseen feature modalities has become increasingly important due to the growing need for cross-disciplinary knowledge integration. A key challenge here is how to align the representation of new modalities with the most relevant parts of the pre-trained model's representation space to enable accurate knowledge transfer. This requires combining feature alignment with target fine-tuning, but uncalibrated combinations can exacerbate misalignment between the source and target feature-label structures and reduce target generalization. Existing work, however, lacks a theoretical understanding of this critical interaction between feature alignment and target fitting. To bridge this gap, we develop a principled framework that establishes a provable generalization bound on the target error, which explains the interaction between feature alignment and target fitting through a novel concept of feature-label distortion. This bound offers actionable insights into how this interaction should be optimized for practical algorithm design. The resulting approach achieves significantly improved performance over state-of-the-art methods across a wide range of benchmark datasets.