Score
Designs and implements image-based re-identification systems that project detections (for example from 3D detections or alternate views) into image space and extract compact appearance embeddings to match identities across frames or cameras. This competence covers building projection modules, decoupling geometric versus appearance modeling in the embedding, and optimizing the embedding extraction and matching pipeline for computational and latency constraints.
Current single-modality person re-identification faces performance bottlenecks under challenging conditions such as low illumination and occlusion. This work systematically reviews the evolution of multimodal person re-identification, encompassing cross-modal tasks including visible-infrared, text-to-image, sketch-based, and non-line-of-sight scenarios. It further introduces, for the first time, a Transformer-based baseline framework for visible-infrared re-identification, which achieves effective matching through cross-modal alignment and modality-invariant feature learning. The proposed comprehensive survey framework and baseline model not only substantially enhance cross-modal retrieval performance but also establish a unified reference benchmark and offer forward-looking directions for future research in the field.
Supervised person re-identification (ReID) suffers from high annotation costs and poor generalizability, hindering large-scale deployment. This paper presents a systematic survey of supervised and unsupervised ReID advancements from 2020 to 2023, proposing a “dual-track comparative” analytical framework. It observes that supervised methods have approached performance saturation—achieving mAP >90% on standard benchmarks—while unsupervised approaches, leveraging clustering-based pseudo-labeling, cross-domain self-supervised pretraining, and attention-enhanced representation learning, have achieved over 25-percentage-point mAP gains on Market-1501 and DukeMTMC-reID, with several methods narrowing the gap to supervised baselines to less than 3%. Crucially, this work provides the first quantitative analysis of convergence trends for both paradigms, identifying key technical pathways—e.g., robust pseudo-label refinement, domain-adaptive contrastive learning, and uncertainty-aware clustering—as essential for transitioning unsupervised ReID toward practical deployment, while highlighting persistent open challenges in scalability, label noise resilience, and cross-scenario generalization.
To address poor robustness and low real-time performance in person detection and cross-camera re-identification (re-ID) under open-world retail and public scenarios—characterized by multi-camera setups, varying illumination, occlusions, and scale changes—this paper proposes a lightweight end-to-end detection-re-ID joint framework. Methodologically, it integrates an enhanced YOLOv8 detector, a lightweight Transformer-based re-ID module, ByteTrack for multi-object tracking, and adaptive feature normalization to jointly optimize detection, localization, and cross-camera matching. Evaluated on real-world retail and street-scene video streams, the framework achieves 92.3% mAP for re-ID, an average per-frame latency of <45 ms, and concurrent processing of 16 HD video streams. Its core contribution lies in the first incorporation of adaptive normalization into an end-to-end joint architecture, significantly improving generalization under complex illumination and occlusion while maintaining high inference efficiency.
This paper addresses the challenge of long-term cross-temporal person re-identification (Re-ID), where significant appearance changes—such as clothing and body shape variations—cause severe performance degradation in cross-camera matching over extended periods. To this end, we introduce CHIRLA, the first large-scale benchmark designed for realistic long-term deployment: it spans seven months, four indoor areas, and seven synchronized cameras, encompassing 22 identities, 5 hours of video, and over one million precisely annotated identity bounding boxes. We propose a controllable long-term appearance variation modeling paradigm, enabled by multi-camera collaborative acquisition, precise temporal synchronization calibration, and a semi-automatic trajectory annotation pipeline. CHIRLA fills a critical gap in evaluating long-term Re-ID systems. Baseline experiments reveal a 47% performance drop in cross-month matching, empirically validating the problem’s difficulty and providing an essential foundation for developing models robust to long-term appearance dynamics.
In person re-identification, conventional class prototypes are fixed at class centroids, limiting discriminative capacity and retrieval performance. To address this, we propose Generalized Class Prototype Selection (GCPS), a method that dynamically learns adjustable class prototypes during both training and inference—replacing static centroids with adaptively selected, highly discriminative embeddings from within each class. GCPS jointly optimizes prototype selection and feature representation to simultaneously enhance inter-class separability and intra-class compactness. Extensive experiments on major benchmarks—including Market-1501, DukeMTMC-reID, and MSMT17—demonstrate consistent superiority over state-of-the-art methods: GCPS achieves absolute improvements of 3.2–5.8% in mAP and 1.9–4.3% in Rank-1 accuracy. These results validate GCPS’s effectiveness and generalizability across diverse re-identification scenarios.
This work addresses the limitations of conventional RGB-based person re-identification, which suffers from privacy leakage, sensitivity to illumination variations, and poor robustness under occlusion. To overcome these challenges, we propose a novel RGB-D multimodal temporal modeling approach that integrates depth images with a temporal Transformer encoder for the first time. This framework effectively fuses RGB and depth information while preserving privacy and employs the Hungarian algorithm to achieve globally optimal cross-view matching. The model is optimized using batch-hard triplet loss and demonstrates competitive performance on the TVPR2, GODPR, and BIWI RGBD-ID datasets: using only the depth modality, it achieves CMC and mAP scores comparable to state-of-the-art methods, thereby validating both its effectiveness and privacy-preserving nature.
本文针对细粒度野生动物重识别问题,提出了一种端到端的检测与重识别模型,采用DINOv2和MegaDescriptor增强潜在查询特征。
This work addresses the challenge of maintaining identity consistency in LiDAR-based 3D multi-pedestrian tracking under occlusion or crowded scenarios, where insufficient integration of appearance cues often leads to frequent identity switches. To tackle this, we propose a lightweight projection framework that decouples geometric and appearance modeling and introduces a cascaded matching strategy for multimodal data association. This approach significantly enhances trajectory continuity in occluded settings while operating under low latency constraints. Experimental results demonstrate that naive linear fusion is highly susceptible to visual noise, resulting in degraded performance, whereas our method achieves a favorable balance between accuracy and robustness on the KITTI pedestrian dataset. The findings validate that the proposed lightweight architecture effectively strikes an optimal trade-off between discriminative power and computational efficiency.
This study addresses the challenges of feature entanglement, cross-modal conflict, and missing modalities in multi-modal re-identification by proposing the Modal framework. This approach innovatively integrates deep unfolding networks with coupled sparse coding to achieve transparent feature disentanglement, while incorporating a text-image differential filtering mechanism to enhance discriminability and robustness. Extensive experiments on four benchmark datasets demonstrate that the proposed framework achieves state-of-the-art performance, effectively mitigating performance degradation under missing modality scenarios. Furthermore, it significantly improves model interpretability, establishing a novel paradigm for multi-modal representation learning that simultaneously ensures transparency and robustness.
This study addresses the challenge of adapting pretrained image matchers to wildlife re-identification, where keypoint annotations are typically unavailable. To this end, we propose the first matcher-level adaptation method based on identity supervision. Building upon a pretrained keypoint correspondence model, our approach performs weakly supervised fine-tuning using only identity labels. By mining positive and negative sample pairs and incorporating contrastive learning to optimize feature correspondences, it achieves efficient domain transfer without requiring geometric ground truth. Experiments demonstrate that the proposed method significantly outperforms both general-purpose matchers and existing state-of-the-art fusion approaches in accuracy on public datasets, while exhibiting strong open-world generalization capabilities.
This study addresses the absence of visual recognition solutions and the challenge of cross-view matching in airport lost luggage scenarios by proposing a DINOv3-based baggage re-identification method. This work pioneers the application of the DINOv3 foundation model to this task, integrating a BNNeck classification head and employing LoRA for parameter-efficient fine-tuning to overcome feature adaptation under few-shot conditions. Experimental results on the MVB benchmark demonstrate that the proposed approach significantly outperforms conventional feature-freezing strategies. These findings effectively validate the generalization capability of vision foundation models under limited data regimes, achieving stable and efficient cross-view baggage retrieval.
This work addresses the challenge of distinguishing appearance-similar individuals in video-based person re-identification by proposing an input-aware, scalable mixture-of-experts architecture. The method employs an input-aware routing mechanism to dynamically select relevant experts and integrates a spatiotemporal feature selection strategy to enhance fine-grained discriminative modeling across both spatial and temporal dimensions. The designed module supports flexible expansion with new experts, thereby improving model generalization. Evaluated on two large-scale video person re-identification benchmarks, the proposed approach significantly outperforms existing state-of-the-art methods, achieving leading performance.