π€ AI Summary
Existing cross-modal person and vessel re-identification methods often overlook modality discrepancies in the frequency domain, limiting their generalization to heterogeneous modality scenarios. To address this, this work proposes the first unified framework that jointly models modality consistency in both spatial and frequency domains: spatial-domain feature distribution alignment is achieved through Gaussian alignment, while identity-aware contrastive learning in the frequency domain ensures discriminative consistency. Designed as a plug-and-play module, the approach seamlessly integrates with diverse heterogeneous modality combinations. Extensive experiments across five datasets and seventeen evaluation protocols demonstrate substantial performance gains over multiple baselines, validating the methodβs generality and effectiveness.
π Abstract
Cross-modal Re-Identification (ReID) aims to retrieve the same identity across heterogeneous imaging modalities and has been widely studied in visible-infrared person ReID and cross-modal ship ReID. Existing methods have achieved promising performance by learning modality consistency in the spatial embedding space, yet often overlook frequency-domain modality discrepancy, particularly in high-frequency representations that are both highly discriminative and modality-sensitive. In addition, most approaches are tailored to specific modality settings, limiting their applicability across diverse cross-modal scenarios. To address these challenges, we propose a Dual-Space Modality Consistency Learning (DSMCL) framework for universal cross-modal ReID. Specifically, DSMCL jointly models spatial feature distribution consistency and frequency-domain discriminative consistency. A Spatial Modality Consistency Learning (SMCL) branch performs Gaussian-based feature alignment, while a Frequency-aware Discriminative Consistency Learning (FDCL) strategy regularizes high-frequency representations through identity-aware cross-modal contrastive learning. By jointly capturing modality-specific characteristics and modality-shared identity cues, DSMCL learns robust representations and establishes a unified framework capable of accommodating diverse heterogeneous modality settings. Moreover, DSMCL is a plug-and-play framework that can be readily integrated into existing cross-modal ReID architectures. Extensive experiments on SYSU-MM01, RegDB, LLCM, HOSS-ReID, and CMShipReID across seventeen evaluation protocols show that DSMCL consistently improves multiple representative baselines.