Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding

📅 2026-02-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing approaches to long video understanding often overlook narrative structure during frame selection, relying solely on query relevance and thereby producing fragmented frames that fail to capture the coherent storyline. To address this limitation, this work proposes a training-free, two-stage frame selection framework. First, wavelet-based multiresolution analysis is employed to detect semantic boundaries by extracting stable semantic change signals from noise, enabling the segmentation of temporally coherent clips. Subsequently, a combination of importance scoring and Maximal Marginal Relevance (MMR) ensures a balance between query relevance and frame diversity. Notably, this is the first study to apply wavelet transforms to semantic boundary detection in videos. The method achieves significant performance gains, improving accuracy by 5.5%, 9.5%, and 6.2% on VideoMME, MLVU, and LongVideoBench, respectively, outperforming current state-of-the-art approaches.

Technology Category

Computer Vision: Video Understanding & Activity AnalysisMachine Learning: Large Multimodal Models (LMMs)Search and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Web query analysis, representation and understandingWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataSemantics and Knowledge: Scalable techniques for the creation, curation, publication, maintenance, and consumption of large, Web-based, structured, reusable, knowledge graphs and ontologies
📝 Abstract
Frame selection is crucial due to high frame redundancy and limited context windows when applying Large Vision-Language Models (LVLMs) to long videos. Current methods typically select frames with high relevance to a given query, resulting in a disjointed set of frames that disregard the narrative structure of video. In this paper, we introduce Wavelet-based Frame Selection by Detecting Semantic Boundary (WFS-SB), a training-free framework that presents a new perspective: effective video understanding hinges not only on high relevance but, more importantly, on capturing semantic shifts - pivotal moments of narrative change that are essential to comprehending the holistic storyline of video. However, direct detection of abrupt changes in the query-frame similarity signal is often unreliable due to high-frequency noise arising from model uncertainty and transient visual variations. To address this, we leverage the wavelet transform, which provides an ideal solution through its multi-resolution analysis in both time and frequency domains. By applying this transform, we decompose the noisy signal into multiple scales and extract a clean semantic change signal from the coarsest scale. We identify the local extrema of this signal as semantic boundaries, which segment the video into coherent clips. Building on this, WFS-SB comprises a two-stage strategy: first, adaptively allocating a frame budget to each clip based on a composite importance score; and second, within each clip, employing the Maximal Marginal Relevance approach to select a diverse yet relevant set of frames. Extensive experiments show that WFS-SB significantly boosts LVLM performance, e.g., improving accuracy by 5.5% on VideoMME, 9.5% on MLVU, and 6.2% on LongVideoBench, consistently outperforming state-of-the-art methods.
Problem

Research questions and friction points this paper is trying to address.

frame selection
long video understanding
semantic boundary
Large Vision-Language Models
narrative structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

wavelet transform
semantic boundary detection
frame selection
long video understanding
training-free
🔎 Similar Papers
Wang Chen
Wang Chen
Individual Researcher
Natural Language ProcessingText GenerationInformation Extraction
Y
Yuhui Zeng
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University
Y
Yongdong Luo
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University
T
Tianyu Xie
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University
L
Luojun Lin
College of Computer and Data Science, Fuzhou University
Jiayi Ji
Jiayi Ji
Rutgers University
Yan Zhang
Yan Zhang
Xiamen University
Statistics
Xiawu Zheng
Xiawu Zheng
Associate Professor, IEEE Senior Member, Xiamen University
Automated Machine LearningNetwork CompressionNeural Architecture SearchAutoML