🤖 AI Summary
Video foundational analysis requires efficient shot boundary detection, sampling pattern identification, and dynamic keyframe extraction. This paper proposes the first unified framework that simultaneously addresses hard-cut and short gradual shot boundary detection, discriminates scan-based sampling modes (progressive, interlaced, and film stretch), and performs motion-adaptive keyframe selection. Our method innovatively integrates motion field estimation with normalized cross-correlation (NCC) features, jointly models inter-frame and intra-frame multi-scale features, and introduces a sparse selective computation strategy—achieving 4× real-time processing without compromising accuracy. Extensive evaluation demonstrates robustness and high precision under challenging conditions including large motion, flash effects, flickering, low contrast, and noise. The framework delivers reliable, efficient foundational analysis to support higher-level video understanding tasks.
📝 Abstract
The detection of shot boundaries (hardcuts and short dissolves), sampling structure (progressive / interlaced / pulldown) and dynamic keyframes in a video are fundamental video analysis tasks which have to be done before any further high-level analysis tasks. We present a novel algorithm which does all these analysis tasks in an unified way, by utilizing a combination of inter-frame and intra-frame measures derived from the motion field and normalized cross correlation. The algorithm runs four times faster than real-time due to sparse and selective calculation of these measures. An initial evaluation furthermore shows that the proposed algorithm is extremely robust even for challenging content showing large camera or object motion, flashlights, flicker or low contrast / noise.