TRIM-ReID: Duplication-Aware Token Reduction and Modality-Aligned Interaction for Multi-Modal Object Re-Identification

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of fine-grained feature loss, interaction noise induced by token redundancy, and high computational overhead in multimodal re-identification. To this end, we propose an efficient cross-modal alignment framework that extracts fine-grained features via dense identity representation learning. We introduce a novel deduplication-aware token reduction mechanism coupled with diversity mining to suppress redundancy. Furthermore, a modality-relation interaction network and a triangular alignment loss are designed to ensure cross-modal semantic consistency under independent token selection. Built upon the DINOv3 encoder, the proposed method achieves state-of-the-art performance on benchmarks such as RGBNT201, significantly enhancing both retrieval accuracy and computational efficiency.
📝 Abstract
Multi-modal object re-identification exploits complementary RGB, near-infrared (NIR), and thermal-infrared (TIR) observations to retrieve target objects. However, existing methods commonly employ visual encoders optimized for global image-text alignment and select tokens using learned importance scores. Such designs fail to preserve fine-grained identity cues or explicitly account for token redundancy, resulting in underrepresented local evidence and duplicated tokens that lead to noisy and costly cross-modal interaction. To address this gap, we propose TRIM-ReID, a compact framework that unifies dense feature extraction, intra-modal token reduction, and inter-modal aligned interaction. Specifically, semantically rich and spatially coherent patch features are extracted by Dense Identity Representation (DIR), which leverages DINOv3 to preserve fine-grained identity information. We then introduce Token Diversity Mining (TDM) to identify complementary local evidence and construct compact modality-specific token sets by suppressing repetitive patches while preserving informative diversity. Retained tokens are subsequently fused by Modal Relational Interaction (MRI) to enable effective information exchange across modalities, while a triangular alignment loss explicitly regularizes their joint relationships to maintain cross-modal semantic consistency under independent token selection. Extensive experiments on RGBNT201, RGBNT100, and MSVR310 demonstrate that TRIM-ReID achieves state-of-the-art performance.
Problem

Research questions and friction points this paper is trying to address.

Multi-modal object re-identification
Token redundancy
Fine-grained identity cues
Cross-modal interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Modal Re-Identification
Token Reduction
Token Diversity Mining
Cross-Modal Alignment
Dense Identity Representation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wanke Xia
Tsinghua University
R
Ruiding Zhu
Anhui University
X
Xingguo Xu
Dalian University of Technology
Zhengbo Zhang
Zhengbo Zhang
Singapore University of Technology and Design
Generative ModelsReinforcement Learning
Dongxia Liu
Dongxia Liu
Tsinghua University
Artificial Intelligence Generated Content
Yuan Jin
Yuan Jin
Apple
Quantum Cascade LasersSemiconductor PhysicsIntegrated Photonics
T
Taojie Zhu
Tsinghua University
Y
Yiting Zhao
Tsinghua University
Y
Yihang Ding
Tsinghua University