Score
Designs and implements methods that map or transform feature representations produced by different modalities or model branches into a common, compatible embedding space so corresponding content aligns across views. This includes building alignment modules, projection heads, and loss functions or distillation procedures (e.g., dual-branch or teacher–student feature distillation) that preserve instance identity and cross-modal correspondence.
This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.
This study investigates the representational potential of foundation models for cross-modal alignment—specifically, whether their unimodal representations inherently capture task-specific semantics and exhibit cross-modal transferability. Method: We formalize “representational potential” and systematically analyze structural regularities and semantic consistency across vision, language, and speech foundation models. Leveraging cross-modal similarity metrics, representation visualization, and neuroscience-inspired evaluation protocols, we assess generalizability and unification capacity across diverse architectures. Contribution/Results: Empirical results demonstrate that pretrained foundation models implicitly acquire semantic invariances requisite for cross-modal alignment—even when trained on unimodal data—thereby exhibiting strong potential as unified multimodal representation backbones. Our work establishes a theoretical framework for cross-modal alignment and introduces a reproducible, multi-faceted evaluation paradigm grounded in representational analysis.
Existing research lacks a systematic analysis of modality–language backbone integration mechanisms in multimodal large language models (MLLMs). Method: We systematically survey 125 MLLMs published between 2021 and 2025, proposing the first large-language-model-centric three-dimensional taxonomy—architectural integration, representation learning, and training paradigms—to unify cross-modal alignment pathways. Leveraging bibliometric analysis, architectural decoupling, and fine-grained modeling of embeddings and loss functions, we characterize evolutionary patterns in fusion granularity, joint representation design, and objective function development. Contribution/Results: This work fills a critical theoretical gap in structured analysis of MLLM fusion mechanisms, identifies shared bottlenecks—including misaligned modality-specific feature hierarchies and suboptimal cross-modal optimization objectives—and delivers a reusable theoretical framework and practical guidelines for developing robust, scalable multimodal foundation models.
This study addresses the challenge that structural heterogeneity in cross-modal feature representations precludes unit-level alignment in conventional knowledge distillation. To overcome this limitation, the authors propose an abstraction mechanism based on vector-quantized codebooks, which transforms teacher features into concept-level anchors to guide student learning. By introducing a task-relevance and compatibility-guided codebook selection algorithm, the proposed method circumvents the assumption of structural consistency across feature spaces, thereby enabling concept-level distillation between heterogeneous modalities without requiring direct alignment. The effectiveness of this framework is validated across multiple classification and semantic segmentation tasks, demonstrating substantial improvements in cross-modal knowledge transfer performance under heterogeneous feature conditions.
Multimodal dataset distillation (MDD) suffers from modality collapse—characterized by excessive intra-modal representation concentration and inter-modal distribution misalignment—exacerbated by asymmetric cross-modal supervision in existing methods, leading to optimization bias. This work first identifies the root cause as an inherent conflict between distillation compression objectives and contrastive learning goals. To address this, we propose RepBlend: a novel framework that weakens overly strong cross-modal constraints via representation blending to enhance intra-modal diversity; introduces modality-specific projection heads and a symmetric projection trajectory matching mechanism to jointly optimize intra-modal diversity and inter-modal alignment. Evaluated on Flickr-30K and MS-COCO under the 100-sample setting, RepBlend achieves +9.4 and +6.3 absolute gains in IR@10 and TR@10, respectively, and accelerates distillation by up to 6.7×, significantly outperforming state-of-the-art methods.
This work addresses a critical limitation in existing independently trained multimodal contrastive models—such as CLIP, SigLIP, and FLAVA—which lack explicit alignment of their representation spaces, particularly in image–text coupling consistency. The authors theoretically demonstrate that embedding spaces from image and text encoders, despite being trained under different architectures and data distributions, can be simultaneously aligned via a single orthogonal mapping. Building on this insight, they propose a unified alignment framework integrating orthogonal mapping modeling, multimodal kernel consistency analysis, and anchor set validation. Notably, the method operates without requiring re-embedding, enabling seamless compatibility with pre-trained models. Extensive experiments across multiple established architectures validate its effectiveness, while also offering a novel perspective on privacy-preserving multimodal representation learning.
Embedding models trained independently on similar data capture stable semantic meanings but yield inconsistent representation spaces, hindering interoperability across models. This work addresses compatibility challenges in multimodal search and model upgrades via orthogonal transformation-based embedding alignment. Theoretically, we derive the first tight Procrustes alignment error bound, proving the existence of a near-isometric orthogonal transformation that approximately preserves pairwise inner products—establishing rigorous theoretical foundations for alignment. Methodologically, we employ efficient Procrustes analysis as a post-hoc alignment procedure, preserving the intrinsic geometric structure of each embedding space while enabling cross-model alignment. Experiments demonstrate substantial improvements in model retraining compatibility, text retrieval fusion accuracy, and cross-modal search performance; our method achieves state-of-the-art results in hybrid multimodal search.
Existing cross-modal contrastive distillation methods overemphasize modality-shared features while neglecting modality-specific information, leading to incomplete 3D representations. To address this, we propose CMCR—a novel framework that systematically identifies, for the first time, the theoretical limitations of contrastive distillation in 3D representation learning. CMCR innovatively integrates geometrically enhanced masked image modeling, voxel occupancy estimation, and a unified multimodal vector-quantized codebook to jointly model both shared and modality-specific features. The unified codebook enables cross-modal semantic alignment, while geometric priors embedded via occupancy modeling enforce structural 3D constraints. Extensive experiments on multiple downstream 3D perception tasks—including 3D object detection and semantic segmentation—demonstrate that CMCR consistently outperforms state-of-the-art image-LiDAR contrastive distillation approaches, substantiating significant improvements in representation completeness and generalization capability.
This study addresses the challenge of unsupervised cross-modal alignment between independently trained models, which typically requires extensive paired data. Grounded in the Platonic Representation Hypothesis, this work proposes a pairing-free alignment method that establishes cross-modal correspondences by exploiting geometric similarities within representation spaces. Specifically, orthogonal mappings are estimated via the Wasserstein Procrustes algorithm combined with geometric initialization techniques to effectively align embedding spaces. Notably, this research provides the first demonstration that coarse-grained cross-modal alignment can be achieved using geometric initialization alone. Extensive experiments across multiple datasets and modalities reveal that the proposed approach significantly outperforms existing baselines under extremely limited paired-data scenarios. Furthermore, its practical utility is validated through successful application to text-to-image generation tasks.
本文提出SPACE方法,利用CLIP的视觉-语言空间中的语义结构解决领域适应问题,通过文本描述作为语义锚点对齐图像的意义而非外观。
This work addresses the high computational cost and difficulty in preserving cross-modal alignment and joint distribution inherent in existing vision-language dataset distillation methods. The authors propose a geometry-aware multimodal distribution matching framework that jointly optimizes at the data, model, and loss levels. Specifically, synthetic samples are initialized via clustering in a shared embedding space, mixed supervision signals are generated through teacher model weight interpolation, and a symmetric contrastive objective leveraging the geometric structure of the unit hypersphere is introduced to align directional features across modalities. This approach substantially reduces distillation overhead while producing compact, semantically faithful, and architecture-agnostic synthetic datasets that effectively retain the original multimodal semantics and alignment performance across multiple text–image retrieval benchmarks.
This work investigates whether the understanding and generation branches of unified multimodal models share a transferable semantic space. To this end, the authors propose a cross-branch semantic guidance framework that extracts intervention-based semantic directions from the understanding branch and transfers them to the generation branch for controllable image synthesis. Their experiments reveal, for the first time, a semantic asymmetry between the two branches: the understanding branch encodes object-level semantics, whereas the generation branch relies more heavily on low-level appearance features. The proposed method enables effective semantic transfer from understanding to generation, significantly improving the semantic fidelity of generated images; however, reverse transfer yields limited gains, demonstrating that architectural unification does not inherently guarantee semantic alignment.
本文提出CrossFeat框架,通过在特征描述子空间学习转换函数,使单模态描述子能够跨模态工作,解决多模态图像匹配问题。