CueKFS: Agentic Cue-Driven Keyframe Selection for Long Video Understanding

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of static keyframes in long-form video question answering, where they often fail to capture implicit information and handle complex queries. We propose a training-free dynamic visual clue retrieval framework that reformulates questions into dynamic clues. By leveraging vision-language models (VLMs) and agent-based reasoning, the method iteratively refines these clues and autonomously re-explores the video. Furthermore, it integrates dynamic prompt decomposition with budget allocation to achieve evidence-driven keyframe selection, thereby overcoming static context constraints. Experimental results demonstrate that our framework achieves state-of-the-art performance across 27 scenarios on three benchmarks, yielding an average improvement of 4.54%. Notably, this approach requires only two VLM invocations while attaining 92% of the maximal similarity gain, highlighting its efficiency and effectiveness for long-video understanding.
📝 Abstract
Keyframe selection (KFS) has long produced compact video summaries for browsing and retrieval, and representative frames for thumbnails. More recently, when conditioned on a question, KFS provides an alternative to uniform sampling for long-video question answering by selecting frames that are more relevant to the question. Most methods rank frames by similarity to the question. Yet a relevant frame may score poorly when the question combines subjects or moments that no single frame shows, or requires implicit information absent from its wording. Other methods try to break down the question into subqueries, but suffer from inaccurate decomposition due to their static initial context. Therefore, we propose CueKFS, a training-free method that reformulates question--frame matching as comparing frames against a set of dynamically generated visual cues. From an initial set of salient frames, we decompose the question into cues. Each cue concurrently probes the video to navigate to its own evidence. A reasoning VLM then agentically revises the cue set against its evidence to re-explore the video. CueKFS then allocates the budget across the surviving cues. Across three benchmarks, CueKFS establishes state-of-the-art results in all 27 evaluated settings with available prior results, achieving budget-averaged gains of up to +4.54% over the previous baseline and a median of only two VLM calls. We further provide a detailed behavioral analysis of CueKFS, showing that agentic cue refinement drives active re-exploration of the video, yielding relative similarity gains of up to 92% over the initial context.
Problem

Research questions and friction points this paper is trying to address.

Keyframe Selection
Long Video Understanding
Video Question Answering
Question Decomposition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Keyframe Selection
Agentic Reasoning
Visual Cues
Training-free
Long Video Understanding
🔎 Similar Papers
No similar papers found.