Semantic Modality Compensation for Unsupervised Visible-Infrared Person Re-identification under Unpaired Settings

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of unpaired settings in unsupervised visible-infrared person re-identification, where the absence of cross-modal identity correspondences leads to feature distortion and pseudo-label noise. We propose a semantic modality compensation framework that reformulates unpaired learning as a semantic compensation problem. Specifically, it decouples identity content from modality style within a shared visual-semantic space and precisely synthesizes missing modal counterparts via prompt composition. Technically, the method integrates augmented dual contrastive learning, CLIP-based prompt learning, and a confidence-gated memory injection mechanism to ensure representation reliability. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods under both paired and unpaired settings, exhibiting particularly significant advantages when severe identity mismatches occur across modalities.
📝 Abstract
Unsupervised visible-infrared person re-identification (USL-VI-ReID) learns person representations that can be compared across modalities without identity annotations. In the unpaired setting, however, identity correspondences between modalities are often incomplete, leaving many identities without an observed counterpart in the other modality. Existing unpaired methods bridge this gap by generating or mapping features for the other modality, mainly by exploiting the statistics of visual features without explicitly separating content that is discriminative for identity from style that is specific to modality. Consequently, the generated features may distort identity cues or inherit bias from the source modality, undermining the reliability of supervision across modalities. We formulate unpaired learning across modalities as a semantic compensation problem and propose Semantic Modality Compensation (SMC), a framework based on prompt composition that decouples identity semantics from modality style within a shared visual semantic space. SMC first constructs a discriminative ReID space through augmented dual contrastive learning, yielding pseudo labels, cluster prototypes, and memory banks for each modality. It then learns visible and infrared modality prompts in the CLIP semantic space and maps clusters obtained from pseudo labels to identity semantic tokens. For each cluster lacking a reliable match in the other modality, SMC combines its identity token with the prompt for the target modality to synthesize a semantic counterpart in the missing modality. The synthesized counterpart is then projected back into the ReID space and injected into a compensation memory through confidence gating. Extensive experiments under both paired and unpaired settings demonstrate that SMC consistently outperforms state-of-the-art methods, with particularly large gains when identity mismatch is severe.
Problem

Research questions and friction points this paper is trying to address.

Unsupervised Visible-Infrared Person Re-identification
Unpaired Settings
Semantic Modality Compensation
Cross-modal Supervision
Identity-Style Decoupling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Modality Compensation
Unsupervised Visible-Infrared Person Re-identification
Prompt Composition
CLIP Semantic Space
Unpaired Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Duanning Chen
Wuhan University
Ke He
Ke He
Univerisity of Canterbury
Machine Learning
B
Bin Yang
Wuhan University
Y
Yongxiang Yao
Wuhan University