Score
Designs and implements algorithms that select a compact set of keyframes from video frame sequences that are both diverse and informative, balancing redundancy reduction and relevance. This includes building DPP-based and iterative diversity samplers, query-conditioned selection, entropy- or self-attention-based scoring (e.g., spectral entropy as a difficulty indicator), and training-free frame-selection heuristics to minimize processed frames while preserving evidential content.
To address the performance degradation of Video-LLMs caused by suboptimal keyframe selection in long videos, this paper proposes MDP3—a training-free, model-agnostic keyframe selection method. MDP3 jointly optimizes for query relevance, list-level diversity, and temporal coherence under a unified framework. It is the first approach to integrate: (i) conditional Gaussian kernels for query-aware frame similarity estimation; (ii) determinantal point processes (DPPs) to model inter-frame diversity; and (iii) a Markov decision process coupled with dynamic programming to enforce temporal contiguity. Theoretical analysis provides a $(1-1/e)$-approximation guarantee for the resulting combinatorial optimization. Extensive experiments on multiple video understanding benchmarks demonstrate that MDP3 significantly outperforms uniform sampling, query-matching heuristics, and other baselines—achieving substantial accuracy gains while maintaining high computational efficiency and strong cross-model generalizability.
Existing keyframe sampling methods struggle to efficiently capture query-relevant evidential cues in long-form video understanding due to limitations in context length and computational cost. This work proposes an evidence-driven keyframe sampling framework grounded in information bottleneck theory, formulating keyframe selection as the maximization of conditional mutual information between selected frames and the query. Through structural decomposition, this objective is transformed into a frame-wise independent scoring problem. To the best of our knowledge, this is the first approach to integrate the information bottleneck principle into query-guided video sampling. By combining a query-conditioned evidence scoring network with a contrastive learning objective, the method substantially improves both training and inference efficiency. Under strict token budgets, it achieves significant performance gains over existing sampling strategies across multiple long-video question-answering benchmarks.
Existing video frame sampling evaluation methods struggle to effectively measure how well sampled frames capture the informativeness and representativeness of a video. To address this limitation, this work proposes STEC, a novel no-reference evaluation metric that integrates spatiotemporal structural entropy with coverage in a lightweight, task-agnostic manner without requiring access to the original reference video. STEC quantifies sampling quality by modeling three key aspects: the spatial information strength of individual frames, the temporal distribution breadth across the sequence, and the non-redundant coverage of visual content. Experiments on MSR-VTT test-1k demonstrate that STEC effectively discriminates among random, uniform, and content-aware sampling strategies and reveals their robustness differences at the individual video level.
This work addresses the high computational cost of video large language models caused by processing dense frames, as well as the limitations of existing keyframe selection methods that often fall into local optima and introduce noisy frames. To this end, the authors propose a training-free frame selection framework that introduces, for the first time, a “directed diversity” metric to unify relevance and diversity into an irreplaceability score. Coupled with a budget-aware adaptive iterative mechanism, the method dynamically optimizes temporal context coverage while preserving core semantics. Evaluated on the LLaVA-Video-7B model across long-video benchmarks, the approach achieves an average performance gain of 12.5%, significantly outperforming baseline strategies such as uniform sampling.
In long-video understanding, visual token overload leads to context overflow, while existing frame-level sparse sampling methods impair motion and event reasoning by disrupting temporal continuity. Method: We propose a key-segment selection paradigm—first systematically demonstrating its superiority over key-frame selection—and introduce explicit temporal coherence modeling. Our approach features F2C, a training-free key-segment sampler, coupled with an adaptive resolution strategy that dynamically balances spatial fidelity and temporal coverage under a fixed token budget; it further integrates short-term coherent segment extraction and efficient token allocation. Contribution/Results: On three major long-video benchmarks—Video-MME, LongVideoBench, and MLVU—our method outperforms uniform sampling by 8.1%, 5.6%, and 10.3%, respectively, significantly enhancing long-range temporal reasoning capability.
Long-form video understanding faces high computational costs and limitations in existing keyframe selection methods, which struggle to adaptively adjust sampling density and coverage based on query content. This work proposes CSES, the first approach to formalize keyframe selection as a coverage optimization problem. Without requiring training, CSES estimates the saliency of frame-query relevance distributions to guide active sampling and constructs a monotone submodular coverage function incorporating semantic relevance, temporal structure, and visual redundancy. A greedy algorithm ensures near-optimal solutions with theoretical guarantees. Evaluated across two benchmarks using four large vision-language models, CSES reduces the number of scored frames by 4–13× and input keyframes by 18.4%–20.5% compared to baselines, while maintaining accuracy and achieving 3.1–5.4× faster frame selection.
Existing approaches to long-form video understanding struggle under the one-shot encoding paradigm to simultaneously achieve broad temporal coverage, fine-grained visual detail, and computational efficiency, often sacrificing critical information through aggressive compression or incurring substantial memory and latency overheads. This work proposes a progressive evidence acquisition framework that generates compact video previews via query-aware adaptive relevance-diversity sampling (AdaRD) and, when model uncertainty arises, triggers a zero-cache on-demand retrieval mechanism to directly fetch high-resolution frames from disk. Requiring no pre-cached frames, the method significantly outperforms current state-of-the-art approaches across seven benchmarks—yielding a 2.59% absolute gain in accuracy on VideoMME, an 8.39% improvement in mIoU on Charades-STA, and a ~33× reduction in visual token consumption.
This work addresses the limitations of existing KV cache compression methods in streaming video understanding, which rely on local heuristics and fail to ensure representative retention of historical visual information. The paper formulates KV cache compression as a coreset selection problem for the first time, proposing a dual-criteria optimization objective in a joint key-value representation space that balances coverage of salient information and numerical diversity. An orthogonality-driven mechanism is introduced to enhance cache quality, and a theoretical connection to log-determinant subset selection is established. Evaluated across four open-source vision-language models and five long-form or streaming video benchmarks, the proposed method significantly outperforms current heuristic-based compression approaches under fixed cache budgets.
This work addresses the challenge of severe temporal redundancy in long video understanding, where dense frame processing is computationally expensive and semantically inefficient. The authors propose a plug-and-play, training-free framework for query- and content-aware keyframe selection that jointly models query relevance and content diversity for the first time. By dynamically allocating a keyframe budget and iteratively selecting frames centered around those most relevant to the query while ensuring semantic richness and diversity, the method can be seamlessly integrated into existing video large language models without fine-tuning. It achieves state-of-the-art performance across multiple long video understanding benchmarks, attaining 67.8% accuracy on LongVideoBench with only 128 frames—surpassing GPT-4o’s 66.7% accuracy using 256 frames.
This study addresses the instability in local event evidence allocation caused by global Top-K competition in sparse video understanding under fixed budgets. To this end, we propose dKFD, a method that partitions budget capacity into early, middle, and late phases based on full-sequence temporal encoding. By integrating a differentiable selector with a phase supervision mechanism, dKFD breaks conventional global competition patterns and establishes an event-centric, controlled structured evidence allocation paradigm. Experimental results demonstrate that dKFD improves Frame AUC by 30.97% on the DoTA dataset, significantly enhancing the alignment between the selector and target events. Furthermore, its effectiveness is validated on the downstream task of Vulnerable Road User (VRU) accident detection.