🤖 AI Summary
Existing approaches to long video understanding often overlook narrative structure during frame selection, relying solely on query relevance and thereby producing fragmented frames that fail to capture the coherent storyline. To address this limitation, this work proposes a training-free, two-stage frame selection framework. First, wavelet-based multiresolution analysis is employed to detect semantic boundaries by extracting stable semantic change signals from noise, enabling the segmentation of temporally coherent clips. Subsequently, a combination of importance scoring and Maximal Marginal Relevance (MMR) ensures a balance between query relevance and frame diversity. Notably, this is the first study to apply wavelet transforms to semantic boundary detection in videos. The method achieves significant performance gains, improving accuracy by 5.5%, 9.5%, and 6.2% on VideoMME, MLVU, and LongVideoBench, respectively, outperforming current state-of-the-art approaches.
📝 Abstract
Frame selection is crucial due to high frame redundancy and limited context windows when applying Large Vision-Language Models (LVLMs) to long videos. Current methods typically select frames with high relevance to a given query, resulting in a disjointed set of frames that disregard the narrative structure of video. In this paper, we introduce Wavelet-based Frame Selection by Detecting Semantic Boundary (WFS-SB), a training-free framework that presents a new perspective: effective video understanding hinges not only on high relevance but, more importantly, on capturing semantic shifts - pivotal moments of narrative change that are essential to comprehending the holistic storyline of video. However, direct detection of abrupt changes in the query-frame similarity signal is often unreliable due to high-frequency noise arising from model uncertainty and transient visual variations. To address this, we leverage the wavelet transform, which provides an ideal solution through its multi-resolution analysis in both time and frequency domains. By applying this transform, we decompose the noisy signal into multiple scales and extract a clean semantic change signal from the coarsest scale. We identify the local extrema of this signal as semantic boundaries, which segment the video into coherent clips. Building on this, WFS-SB comprises a two-stage strategy: first, adaptively allocating a frame budget to each clip based on a composite importance score; and second, within each clip, employing the Maximal Marginal Relevance approach to select a diverse yet relevant set of frames. Extensive experiments show that WFS-SB significantly boosts LVLM performance, e.g., improving accuracy by 5.5% on VideoMME, 9.5% on MLVU, and 6.2% on LongVideoBench, consistently outperforming state-of-the-art methods.