I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning

πŸ“… 2026-08-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the limitations of existing video reasoning methods, which are largely confined to simplified video–text settings and struggle to support identity-aware cross-modal association, action understanding, and temporal reasoning. To bridge this gap, we introduce the Identity-Conditioned Querying (ICQ) task, which leverages both video content and reference character images to enable identity grounding, action recognition, and temporal inference. We present ISYV-Bench, a six-level difficulty evaluation benchmark, along with ISYV-75K, a 75K-scale training set, constructed via an automated labeling pipeline incorporating multi-stage validation and human review. Furthermore, we propose a key clip modeling framework that operates without shot boundary annotations and design a multimodal large language model tailored for ICQ. Experiments show that current open- and closed-source models perform modestly on ISYV-Bench, whereas our ISYV-Model significantly outperforms strong baselines and approaches the performance of leading closed-source systems on several metrics.
πŸ“ Abstract
Real-world video reasoning often involves multimodal, multi-source inputs, whereas existing video reasoning tasks typically assume a simplified video-text setting, limiting identity matching and person-centric reasoning. To bridge this gap, we introduce the Identity-conditioned Queries (ICQ) task, in which models are required to jointly associate and interpret an input video and a reference image of a person, and leverage this conditioning to address identity grounding, behavior understanding, and temporal reasoning, among other challenges. Building on ICQ, we present ISYV (I Seek You in Videos), a systematic solution comprising three components: (1) ISYV-Bench, a challenging evaluation benchmark with 1,377 real-world complex videos and 1,377 question-answer pairs, organized into six difficulty levels spanning capabilities from identity recognition to causal reasoning; (2) ISYV-75K, a large-scale training set of 75K high-quality samples constructed via automated annotation, multi-stage verification, and manual review; and (3) ISYV-Framework, containing an ICQ-oriented model and training strategy for learning to exploit informative video shots without additional shot-level annotations. Extensive experiments show that both mainstream closed-source and open-source MLLMs struggle on ISYV-Bench, especially in cross-domain identity matching and long-horizon tracking. ISYV-Model outperforms strong baselines and in some aspects approaches closed-source performance. Overall, ISYV provides a unified task definition, scalable datasets/benchmarks, and modeling insights for person-centric video reasoning.
Problem

Research questions and friction points this paper is trying to address.

person-centric video reasoning
identity grounding
multimodal reasoning
video understanding
identity matching
Innovation

Methods, ideas, or system contributions that make the work stand out.

Identity-conditioned Queries
Person-centric Video Reasoning
Video Benchmark
Multimodal Learning
Temporal Reasoning
S
Shibo Gao
Beijing Jiaotong University; HUJING Digital Media & Entertainment Group; MAIS Institute of Automation, Chinese Academy of Sciences
C
Chongxiao Wang
MAIS Institute of Automation, Chinese Academy of Sciences
C
Chenglong Huang
MAIS Institute of Automation, Chinese Academy of Sciences
J
Jie Ma
MAIS Institute of Automation, Chinese Academy of Sciences
Haolin Shi
Haolin Shi
University of Science and Technology of China
3D AIGCComputer Vision
Fei Ding
Fei Ding
Unknown affiliation
J
Jing Li
HUJING Digital Media & Entertainment Group
Q
Qiang Lyu
School of Computer Science and Technology, University of Chinese Academy of Sciences
Yangyang Liu
Yangyang Liu
casia
OCRDeep Learning
Yang Liu
Yang Liu
AP, Tongji University; Ph.D., Fudan University & University of Toronto; B.E., Nanjing University
Signal processingComputer visionComputing
Jun Liu
Jun Liu
Professor, Lancaster University
Computer VisionDigital HealthMachine LearningHuman Behavior Analysis
L
Linlin Huang
Beijing Jiaotong University
P
Peipei Yang
MAIS Institute of Automation, Chinese Academy of Sciences