Score
Techniques for aligning multi-modal or repeated images (handling motion, modality-specific artifacts, and missing data) to produce consistent, registered targets suitable for training and analysis across scans and instruments.
This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.
This study addresses the limited robustness of anatomical landmark matching across multi-modal (CT/MRI) and longitudinal medical images. We propose a training-free, deep learning–free consistency-guided point matching method grounded in classical point-set registration frameworks. Our approach introduces cross-image consistency constraints to enhance geometric coherence, enabling efficient, CPU-based execution with tunable accuracy–speed trade-offs. Unlike existing supervised or large-scale annotated-data–dependent methods, our framework achieves significant improvements in landmark localization accuracy and generalizability on the Deep Lesion Tracking benchmark and multiple internal/public longitudinal datasets. To our knowledge, it is the first unsupervised method to deliver high-stability cross-modal anatomical correspondence. The resulting solution is lightweight, clinically deployable, and plug-and-play—offering a reliable alternative for image-guided navigation in real-world clinical settings.
Deformable registration between highly heterogeneous modalities—e.g., PET/FA and MRI/CT—is challenging for conventional unsupervised methods, as cross-modality similarity metrics (e.g., NCC, MI) fail, leading to severe deformation artifacts. Method: We propose M2M-Reg, the first “multi-to-single” supervised paradigm: a multi-modal registration model is trained exclusively using single-modality similarity losses (e.g., NCC or MI on MRI–MRI pairs), eliminating the need for handcrafted cross-modality metrics, ground-truth deformation fields, or segmentation labels. To ensure diffeomorphism, we introduce GradCyCon—a novel gradient-based cyclic consistency regularizer—and integrate it with a differentiable cycle-consistent registration network and a diffeomorphic flow generation architecture. Results: On the ADNI dataset, M2M-Reg achieves state-of-the-art Dice scores for PET–MRI and FA–MRI registration, improving over prior methods by up to 2×, while markedly suppressing image distortion and preserving anatomical plausibility.
This study systematically investigates the impact of preprocessing—particularly color normalization—on cross-modal registration accuracy between H&E-stained digital pathology images and nonlinear multimodal images. Four color normalization methods (CycleGAN, Macenko, Reinhard, and Vahadane) are evaluated, each combined with enhancement strategies including intensity inversion, contrast adjustment, intensity normalization, and denoising; rigid and non-rigid multi-resolution registration is performed using VALIS. Validation on 20 real tissue samples demonstrates that CycleGAN-based preprocessing yields the lowest median and average relative target registration error (MMrTRE/AMrTRE), with point-based evaluation further confirming its superior performance. To our knowledge, this is the first work to establish CycleGAN’s distinct advantage in histopathological cross-modal registration, empirically validating color transformation as a critical preprocessing step for enhancing registration robustness. The findings provide a reproducible, optimized pipeline for spatial biomarker localization and multimodal tissue reconstruction.
In medical object detection, joint training on multimodal images (e.g., X-ray, CT, MRI) suffers from degraded performance due to inter-modal statistical heterogeneity and discontinuous query representation spaces. To address this, we propose QueryREPA—a framework that achieves cross-modal query alignment without modifying the architecture of DETR-based detectors. Our approach comprises three key components: (1) lightweight modality tokens derived from textual descriptions; (2) a Multimodal Contextual Attention (MoCA) mechanism to enhance cross-modal query interaction; and (3) a contrastive learning–based pre-alignment strategy for query representations. QueryREPA significantly improves mean Average Precision (mAP) under multimodal joint training, incurs negligible computational overhead, requires no additional annotations, and effectively enhances model generalization and cross-modal consistency.
Multimodal medical imaging suffers from misalignment across modalities due to the absence of paired registration data between arbitrary modality pairs. Method: This paper proposes M³Bind, the first framework to achieve joint multimodal alignment without explicit inter-modal pairing by leveraging text as a shared semantic mediator. Built upon the CLIP architecture, it introduces modality-specific textual space fine-tuning and knowledge distillation to construct a unified, shared text encoder—preserving each modality’s original image–text alignment capability while enabling zero-shot and few-shot cross-modal retrieval and classification. Results: Extensive evaluation across X-ray, CT, fundus, ECG, and histopathology images demonstrates state-of-the-art performance across multiple tasks, with significant improvements in zero-shot classification accuracy and cross-modal retrieval recall.
This work addresses the significant modality gap between medical imaging and clinical text in shared representation spaces, which leads to insufficient semantic alignment and hampers cross-modal retrieval and understanding performance. To tackle this challenge, the authors propose a modality-agnostic contrastive learning framework that systematically mitigates modality discrepancies in medical settings through optimized embedding space geometry and joint modeling strategies. This approach overcomes the limitations of conventional CLIP-based methods in medical domains and achieves, for the first time, a general and efficient semantic alignment between medical images and clinical text. Experimental results demonstrate substantial improvements in both cross-modal retrieval accuracy on radiology images paired with clinical reports and the quality of generated image captions.
This work addresses the challenges in multimodal image registration, where modality-specific information often leaks into the shared feature space and existing methods struggle to jointly model global rigid alignment and local non-rigid deformations. To overcome these limitations, the authors propose HRNet, which employs a shared backbone enhanced with modality-specific batch normalization (MSBN) and introduces a cross-scale decoupling and adaptive projection module (CDAP) to effectively suppress modality interference. Furthermore, a hybrid parameter prediction module (HPPM) is designed to unify the prediction of rigid transformations and non-rigid deformation fields within an end-to-end, non-iterative framework. The proposed method achieves state-of-the-art performance in both rigid and non-rigid registration across four multimodal datasets.
This work addresses the challenges of modality missing and temporal misalignment in multimodal data caused by sensor failures, asynchronous sampling, and network delays, as well as limitations of existing approaches in inaccurate sample-level alignment and class imbalance. To tackle these issues, the paper proposes a penalty-based many-to-many alignment clustering model grounded in a dual learning mechanism. The method integrates semantic and structural priors from each modality to enhance cross-modal consistency at both local and global levels, while a penalty mechanism refines alignment accuracy and mitigates excessive data concentration. Experimental results demonstrate that the proposed approach significantly improves alignment and clustering performance on incomplete and temporally disordered multimodal data.
Existing vision-language models exhibit limited performance on medical image–text tasks and lack effective tools to quantify inter-modal information imbalance. This work proposes the Asymmetric Spectral Alignment Score (SAS), introducing for the first time a directional alignment metric that projects multimodal representations onto the principal component basis of an anchor modality and computes modality-wise correlations weighted by eigenvalues. SAS reveals an asymmetry in which medical images retain richer structural information than clinical text. Integrated into an evaluation framework encompassing 15 vision-language models and six alignment metrics, SAS demonstrates the strongest correlation with bidirectional retrieval performance under label-free conditions, offering a practical and interpretable tool for assessing medical multimodal models.
This work addresses the challenge of intraoperative liver tumor segmentation in CT, where tumors exhibit extremely low contrast against surrounding tissue, rendering them nearly invisible. In contrast, preoperative MRI clearly delineates lesions. The authors propose the first end-to-end framework for cross-modal registration and weakly supervised segmentation that operates under the extreme setting where pathological structures are entirely absent in the target modality (CT). By leveraging MRI-to-CT registration to generate pseudo-labels, the method enables segmentation of otherwise invisible tumors. It integrates MSCGUNet for multimodal registration and UNet for segmentation, explicitly revealing two core challenges: domain shift and feature absence. Experiments show a Dice score of 0.72 on the CHAOS dataset for healthy liver segmentation, but performance drops sharply to 0.16 on real clinical data containing tumors, highlighting the fundamental limitations of current weakly supervised approaches for truly invisible pathology segmentation.