🤖 AI Summary
This study addresses the challenges of fine-grained feature loss, interaction noise induced by token redundancy, and high computational overhead in multimodal re-identification. To this end, we propose an efficient cross-modal alignment framework that extracts fine-grained features via dense identity representation learning. We introduce a novel deduplication-aware token reduction mechanism coupled with diversity mining to suppress redundancy. Furthermore, a modality-relation interaction network and a triangular alignment loss are designed to ensure cross-modal semantic consistency under independent token selection. Built upon the DINOv3 encoder, the proposed method achieves state-of-the-art performance on benchmarks such as RGBNT201, significantly enhancing both retrieval accuracy and computational efficiency.
📝 Abstract
Multi-modal object re-identification exploits complementary RGB, near-infrared (NIR), and thermal-infrared (TIR) observations to retrieve target objects. However, existing methods commonly employ visual encoders optimized for global image-text alignment and select tokens using learned importance scores. Such designs fail to preserve fine-grained identity cues or explicitly account for token redundancy, resulting in underrepresented local evidence and duplicated tokens that lead to noisy and costly cross-modal interaction. To address this gap, we propose TRIM-ReID, a compact framework that unifies dense feature extraction, intra-modal token reduction, and inter-modal aligned interaction. Specifically, semantically rich and spatially coherent patch features are extracted by Dense Identity Representation (DIR), which leverages DINOv3 to preserve fine-grained identity information. We then introduce Token Diversity Mining (TDM) to identify complementary local evidence and construct compact modality-specific token sets by suppressing repetitive patches while preserving informative diversity. Retained tokens are subsequently fused by Modal Relational Interaction (MRI) to enable effective information exchange across modalities, while a triangular alignment loss explicitly regularizes their joint relationships to maintain cross-modal semantic consistency under independent token selection. Extensive experiments on RGBNT201, RGBNT100, and MSVR310 demonstrate that TRIM-ReID achieves state-of-the-art performance.