Dual-Edged Homogeneous-Modality Similarity: Towards Visible-Infrared Modality-Incomplete Person Re-Identification with Modality Adaptive Matching

๐Ÿ“… 2026-07-21
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenges of visible-infrared person re-identification in open-world scenarios, where uncertain modality combinations and intra-modality interference lead to matching conflicts and degraded robustness. To this end, it formally introduces a novel taskโ€”Visible-Infrared Modality-Incomplete Person Re-Identification (VIMI-ReID)โ€”and establishes two benchmark datasets, SYSU-VIMI and RegDB-VIMI. The authors propose a Modality-Adaptive Matching Transformer (MAMT) featuring a dual-path architecture that disentangles modality-specific and shared features. MAMT integrates a Divergence Transformer, a Shared Transformer, and a modality-adaptive matching module, trained end-to-end with a divergence loss. Experimental results demonstrate that MAMT significantly outperforms existing methods across arbitrary modality combinations, confirming its effectiveness and adaptability under complex modality conditions.
๐Ÿ“ Abstract
Visible-Infrared Person Re-Identification (VI-ReID) operates under a closed-world assumption, where queries and galleries are from heterogeneous modalities. However, in open-world scenarios, both sets are likely to contain homogeneous and heterogeneous modality images. A query may consist of visible-only, infrared-only, or mixed-modality images, while galleries present multi-modal images over long-term collection. Under these conditions, VI-ReID methods, built on a heterogeneous-modality retrieval paradigm, suffer from three trustworthiness challenges: matching conflicts due to high homogeneous-modality similarity, interference from modality uncertainty, and robustness degradation induced by unknown modality combinations. They fail to meet the requirements of trustworthy visual recognition in reliability, consistency, and dynamic adaptability. To address these challenges, we formalize the Visible-Infrared Modality-Incomplete Re-Identification (VIMI-ReID) task. We reorganize existing datasets to construct the SYSU-VIMI and RegDB-VIMI benchmarks. The unpredictable modality combinations and inherent similarity of homogeneous-modality samples in VIMI-ReID cause a significant performance drop in existing VI-ReID methods. We propose the Modality Adaptive Matching Transformer (MAMT). It employs a Divergence Transformer Module (DTM) and a Shared Transformer Module (STM) to extract modality-specific and modality-shared features, respectively. Guided by a divergence loss, the DTM enriches modality-specific features with modality-style information to enhance discriminability within the same modality. A Modality Adaptive Matching Module (MAM) dynamically fuses features according to the query-gallery modality relationship, enabling stable matching under arbitrary and uncertain modality conditions. Extensive experiments on the VIMI benchmarks demonstrate the effectiveness and adaptability of MAMT.
Problem

Research questions and friction points this paper is trying to address.

Visible-Infrared Person Re-Identification
Modality-Incomplete
Homogeneous-Modality Similarity
Open-World Scenario
Trustworthy Recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Modality-Incomplete Re-Identification
Modality Adaptive Matching
Divergence Transformer
Visible-Infrared Person Re-Identification
Dynamic Feature Fusion