RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the dual gap of query comprehension and evidence interpretation in frame selection for long video understanding. To this end, it proposes a training-free reflective agent framework that decouples query interpretation from evidence retrieval. Specifically, the method leverages a lightweight video large language model to reconstruct sub-queries and integrates an embedding model for precise evidence localization, achieving closed-loop optimization through a novel coupled reflection mechanism. Experimental results demonstrate that the proposed framework significantly enhances long video understanding performance across multiple benchmarks. These findings further validate the effectiveness of integrating weak components to augment strong models, offering a promising paradigm for efficient and accurate long-form video reasoning without additional training overhead.
📝 Abstract
Video large language models (Vid-LLMs) excel at diverse video-language tasks by reasoning over selected frames. However, frame selection for long videos remains challenging, as it requires retrieving relevant frames distributed across segments from a large candidate pool given complex queries. This paper investigates dominant approaches to long-video frame selection from a task-decomposition perspective, identifying two key challenges: the Query Comprehension Gap in similarity-based methods and the Interpretation--Selection Gap in judgment-based methods. To address them, we propose RACER, a training-free reflective agentic framework that decomposes long-video frame selection into query interpretation driven by a lightweight Vid-LLM and evidence localization supported by an embedding model serving as a retrieval tool. Specifically, the Vid-LLM is responsible solely for reformulating the complex query into sub-queries that make implicit information requirements explicit, mitigating the Query Comprehension Gap. Meanwhile, the retrieval tool leverages these sub-queries to localize relevant evidence, relieving the Vid-LLM of direct frame selection and thus addressing the Interpretation--Selection Gap. Finally, the retrieved frames are fed back to the Vid-LLM for sub-query refinement, forming a reflection loop that iteratively improves query interpretation and frame selection. Experiments across multiple benchmarks show that RACER consistently improves long video understanding. Notably, RACER achieves effective frame selection even with limited-capability components, demonstrating that agentic integration enables these components to enhance more capable Vid-LLMs.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reflective Agentic Framework
Long Video Understanding
Frame Selection
Query Interpretation
Tool-Based Retrieval
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yiyang Huang
Department of Electrical and Computer Engineering, Northeastern University
Yitian Zhang
Yitian Zhang
Northeastern University
computer vision
Yizhou Wang
Yizhou Wang
Northeastern University
AILLMComputer VisionAnomaly Detection
Jianglin Lu
Jianglin Lu
Northeastern University
Machine Learning
Q
Qihua Dong
Department of Electrical and Computer Engineering, Northeastern University
H
Hailing Wang
Department of Electrical and Computer Engineering, Northeastern University
Huimin Zeng
Huimin Zeng
Northeastern University
computer vision
Mingyuan Zhang
Mingyuan Zhang
Northeastern University
Sparse NetworkLLM Post-trainingMulti-modal LearningMulti-view Learning
Y
Yun Fu
Department of Electrical and Computer Engineering, Northeastern University; Khoury College of Computer Science, Northeastern University