Score
Design and evaluate algorithms that align representations of corresponding samples across two or more views or modalities by constructing and optimizing sample-level matching objectives (e.g., contrastive or CLIP-style losses). This includes building pairing strategies and loss terms that support one-to-many or incomplete pairings, discourage collapse or over-aggregation, and compensate for class-size or sample-imbalance during fusion.
This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.
Conventional multimodal learning assumes explicit representation alignment is universally beneficial, yet its causal impact remains unverified. Method: This work systematically investigates the conditional effects of enforced alignment, proposing a controllable contrastive learning module that dynamically modulates alignment strength and establishes a quantitative relationship between alignment strength and inter-modal information redundancy—derived via information decomposition and synthetic data modeling. Contribution/Results: We demonstrate that alignment efficacy is not universal: strong alignment improves performance under high redundancy but harms modality-specific representation learning under low redundancy. To address this, we introduce a balanced mechanism that jointly preserves shared semantics and modality-specific characteristics. Empirical evaluation on both synthetic and real-world benchmarks confirms that this mechanism significantly enhances the generalization capability of unimodal encoders. Our core contribution is the formal establishment of “redundancy-dependent alignment”—a principled, interpretable, and tunable paradigm for multimodal representation learning.
Existing approaches struggle to model high-order dependencies among more than two modalities and lack a unified principle for balancing information retention and compression. This work introduces the information bottleneck principle into arbitrary multimodal alignment for the first time, proposing a One-vs-All multimodal alignment framework. By optimizing each modality’s sufficiency and minimality with respect to all others, the method derives a computable contrastive lower bound and a minimality regularizer. It further integrates parameter-free geometry-aware projection and a distribution-dependent upper-bound regularizer to effectively capture high-order interactions and geometric structures. The proposed approach achieves consistently strong and state-of-the-art performance across diverse tasks, including classification, regression, modality-agnostic evaluation, and cross-modal retrieval.
Existing contrastive learning methods define similarity exclusively over semantically consistent sample pairs, overlooking latent structural similarities inherent in semantically dissimilar pairs. To address this limitation, we propose SimLAP—a novel framework that redefines positive pairs not as semantically identical samples but as learnable, discriminative subspace-aligned pairs. SimLAP jointly optimizes pairwise similarity estimation and subspace projection within an end-to-end training paradigm, incorporating both contrastive loss and explicit subspace alignment constraints. By uncovering and leveraging structural similarities among inter-class samples residing in shared latent subspaces, SimLAP breaks from conventional similarity modeling paradigms. Extensive experiments across multiple benchmarks demonstrate its effectiveness: SimLAP significantly improves few-shot transfer performance and model robustness under distribution shifts. Moreover, it offers a new perspective on unsupervised similarity learning—shifting focus from semantic identity to geometric consistency in learned representation subspaces.
This paper identifies the “reference mismatch” problem in text-to-image diffusion model alignment—where preference-based methods like DPO suffer substantial performance degradation when the reference model diverges from the target model. To address this, we propose **Margin-aware Preference Optimization (MPO)**, a reference-free framework that models pairwise preferences via the Bradley–Terry model and directly optimizes the likelihood margin between preferred and dispreferred samples, eliminating reliance on a reference model altogether. MPO introduces the first margin-optimization paradigm for diffusion alignment, with gains amplifying as reference mismatch intensifies. Experiments demonstrate that MPO consistently outperforms DPO and DreamBooth across five key alignment tasks: safe generation, style transfer, cultural representation, personalization, and general-purpose alignment. Moreover, MPO achieves 15% faster training and reduced GPU memory consumption.
This work addresses the modality gap in multimodal representation learning induced by the InfoNCE objective, which manifests as a conflict between inter-modal alignment and uniformity, as well as intra-modal alignment inconsistencies. The paper proposes the first framework that decouples alignment and uniformity in multimodal learning, employing Hölder divergence–based alignment optimization alongside a dedicated uniformity loss to effectively mitigate these conflicts. Theoretically, the proposed objective is shown to serve as a valid proxy for the global Hölder divergence between multimodal distributions. Notably, the method requires no task-specific components and consistently improves performance across both discriminative tasks (e.g., retrieval) and generative tasks (e.g., UnCLIP), demonstrating its generality and effectiveness.
Concept alignment lacks a unified definition, and existing methods optimize divergent objectives under the same terminology, obscuring its fundamental nature. This work formalizes its multidimensional structure by decomposing it along two axes—“alignment targets” and “alignment levels”—and identifies four distinct alignment properties, revealing that current approaches satisfy only subsets of these. To address this limitation, we propose Coupled Sparse Autoencoders (CoSAE), a framework that jointly optimizes multiple alignment objectives, alongside InterVenchA, an interventional evaluation benchmark. Experiments demonstrate that optimizing a single objective fails to reliably recover other alignment properties, whereas CoSAE achieves strong instance-level conceptual consistency using merely 0.1% paired data.
This work addresses the unclear interaction between feature alignment and target fitting in cross-modal fine-tuning, which often leads to a mismatch between feature-label structures across source and target domains, thereby degrading generalization. For the first time, this study theoretically characterizes their relationship by introducing the notion of “feature-label distortion,” and establishes a provable generalization bound on target error. Based on this analysis, a principle for joint optimization of alignment and fitting is derived. The resulting framework offers interpretable and actionable design guidelines for cross-modal fine-tuning. Extensive experiments demonstrate that the proposed method significantly outperforms current state-of-the-art approaches across multiple benchmark datasets, confirming its effectiveness and broad applicability.
This work addresses the lack of principled understanding in existing literature regarding the choice between cross-attention and feature concatenation strategies for multimodal fusion, which has largely relied on empirical heuristics. Through controlled experiments and theoretical analysis, we demonstrate for the first time that feature alignment quality is the key determinant of fusion strategy performance: under pre-aligned features, concatenation consistently outperforms cross-attention by 4.1–5.1 percentage points across all data scales, with its advantage becoming more pronounced as alignment degrades. Building on this insight, we develop a theoretical decision framework grounded in sample complexity and validate our findings using features extracted from ResNet-18 and CLIP ViT-B/32 on controlled datasets.
This study addresses the limited generalizability of existing matching solvers, which struggle to uniformly handle a broad spectrum of matching problems. To overcome this, we propose a unified framework grounded in duality theory that, for the first time, leverages difference-of-convex (DC) decomposition to reformulate quadratic matching objectives, such as Gromov-Wasserstein distances, into implicit registration problems. Coupled with a modular algorithmic design, this framework efficiently supports the alignment and processing of multimodal data, including graphs and point clouds. Our contributions provide rigorous convergence guarantees for quadratic matching while extending its applicability to novel scenarios such as fracture matching. By enabling efficient optimization on large-scale datasets, this work significantly broadens the scope of tractable matching problems.
本文探讨了减少CLIP中图像-文本模态差距并不总能提高零样本分类准确率的问题,通过分析决策结构提出预测级集中现象,并建议评估方法应考虑对下游预测结构的影响。