Achieving Text-based Person Retrieval with Any Granularity

๐Ÿ“… 2026-07-23
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge of text-based person retrieval in real-world scenarios, where textual descriptions exhibit uncertain and varying levels of granularity. To tackle this issue, the authors propose a novel paradigm that supports arbitrary description granularities. Their key contributions include constructing UFine6926-MGโ€”the first dataset annotated across a five-level granularity hierarchyโ€”and introducing MG-Eval, a many-to-many evaluation benchmark that better reflects real-world semantic complexity. Furthermore, they present the CMAM framework, which achieves cross-modal alignment across multiple granularities through orthogonal expert perception, probabilistic alignment, and granularity-consistent reasoning. Experimental results demonstrate that CMAM significantly outperforms existing methods across all granularity levels, establishing the first robust baseline and a systematic evaluation protocol for multi-granularity text-based person retrieval.
๐Ÿ“ Abstract
Text-based person retrieval faces a critical but under-explored challenge: the inherent uncertainty of query granularity in real-world scenarios. This paper introduces a new paradigm, Text-based Person Retrieval with Any Granularity, and provides a systematic solution. First, we formalize a five-level granularity spectrum and construct UFine6926-MG, a high-quality multi-grained dataset annotated comprehensively at all granularities via a novel Multi-grained Text Annotation Engine. Second, acknowledging that coarse queries naturally correspond to multiple valid candidates, we propose MG-Eval, a holistic evaluation benchmark with progressively detailed texts and cross-identity labels that reflect real-world semantics, alongside tailored evaluation metrics and protocols. Third, after a comprehensive diagnosis reveals the systemic limitations of existing research, we propose the Cross-modal Multi-grained Aligning and Matching (CMAM) framework. CMAM achieves granularity-aware retrieval through: 1) orthogonal-expert perception to disentangle granularity-specific features; 2) probabilistic alignment to model many-to-many matches under query uncertainty; and 3) granularity-consistent reasoning to steer feature learning via joint cross-modal granularity verification. Experiments demonstrate that CMAM significantly outperforms state-of-the-art methods across all granularity levels. This work establishes a foundational benchmark and a robust baseline, paving the way for more practical person retrieval systems.
Problem

Research questions and friction points this paper is trying to address.

text-based person retrieval
query granularity
granularity uncertainty
multi-grained retrieval
cross-modal matching
Innovation

Methods, ideas, or system contributions that make the work stand out.

text-based person retrieval
multi-grained representation
cross-modal alignment
granularity-aware retrieval
probabilistic matching
๐Ÿ”Ž Similar Papers
No similar papers found.
Jialong Zuo
Jialong Zuo
Zhejiang University
Speech SynthesisVoice Conversion
Hanyu Zhou
Hanyu Zhou
School of Computing, National University of Singapore
Scene UnderstandingMultimodal LearningEvent CameraDomain Adaptation.
D
Dongyue Wu
National Key Laboratory of Multispectral Information Intelligent Processing Technology, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China
Y
Yongtai Deng
National Key Laboratory of Multispectral Information Intelligent Processing Technology, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China
M
Mengdan Tan
National Key Laboratory of Multispectral Information Intelligent Processing Technology, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China
Nong Sang
Nong Sang
Huazhong University of Science and Technology
Computer Vision and Pattern Recognition
C
Changxin Gao
National Key Laboratory of Multispectral Information Intelligent Processing Technology, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China
Xiang Bai
Xiang Bai
Huazhong University of Science and Technology (HUST)
Computer VisionOCR