Score
Designs algorithms or pipelines that select compact sets of representative frames (keyframes) from videos, with selection optimized for a specific downstream task or constraint. Implementations can be training-free or trained and often exploit cues such as scene dynamics, scene-awareness, or JPEG file-size/encoding artifacts to identify informative frames for downstream analysis, agent guidance, or procedural learning.
This work addresses the challenge of fine-grained, frame-level controllable generation in video diffusion models. We propose a training-free, inference-only frame-level guidance framework that requires neither model training nor fine-tuning. Our method introduces a novel training-free latent-space mechanism for per-frame signal injection, coupled with lightweight latent reweighting and cross-frame consistency optimization—enabling precise control over keyframes, sketches, style reference images, and depth maps, all without modifying model parameters. Crucially, we achieve this via gradient-free latent iteration, balancing control fidelity with global temporal coherence. Experiments demonstrate high-quality, controllable video generation across keyframe-guided synthesis, style transfer, and video looping tasks. The framework is fully compatible with arbitrary pre-trained video diffusion models, incurs zero training overhead, and significantly lowers deployment barriers.
To address the lack of generality in scene segmentation and keyframe extraction across heterogeneous video domains—spanning short videos, films, archival footage, and surveillance streams—this paper proposes a dynamic adaptive scene segmentation framework and a lightweight, model-free keyframe scoring mechanism. The former achieves robust granularity control via video-length-driven threshold switching, hybrid segmentation strategies, and interval-based partitioning; the latter introduces an interpretable composite frame quality metric integrating sharpness, brightness, and temporal distribution. Neither component requires pre-trained models, ensuring high accuracy, computational efficiency, and transparency. The method has been deployed in a commercial video analytics platform, enabling high-throughput operation across media, education, research, and security applications. Empirical results demonstrate significant improvements in UI preview generation, semantic embedding construction, and content filtering—enhancing both accuracy and scalability.
To address the challenges of limited context windows and semantic incoherence in keyframe selection for long-video understanding, this paper proposes a scene-driven K-frames method. It leverages semantic-aware segment localization to generate query-relevant, temporally contiguous, and multi-scale adjustable keyframe sequences. Our approach is the first to achieve scene-aware frame selection without requiring additional annotations and supports arbitrary keyframe counts. We introduce the PeakClips dataset and design a three-stage progressive training paradigm that jointly incorporates supervised fine-tuning and reinforcement learning to optimize downstream task performance. Extensive experiments on mainstream long-video understanding benchmarks demonstrate significant improvements over state-of-the-art methods. The proposed framework delivers an efficient, interpretable, and plug-and-play solution for multi-scale keyframe selection.
In unsupervised video object segmentation, models often suffer from unstable predictions due to over-reliance on motion cues such as optical flow. To address this, we propose the “Motion-as-Option” mechanism, which decouples motion information into an optional module: during training, optical flow inputs to the motion encoder are randomly replaced with RGB frames, and an adaptive output selection algorithm dynamically fuses predictions from parallel motion and appearance pathways. This is the first approach to enable non-mandatory modeling of motion cues, integrating motion-appearance collaborative representation learning with stochastic input masking. Our method achieves state-of-the-art performance on DAVIS and FBMS benchmarks, demonstrating significantly improved robustness against anomalous motion disturbances and yielding a 23% gain in prediction stability.
Creating 2D animations is labor-intensive, requiring precise coordination of multiple visual elements across time—a significant barrier for designers at all experience levels. To address this, we introduce Keyframer, a novel tool enabling natural language–driven generation and interactive editing of SVG-based animations. Methodologically, we propose a *decompositional prompting* strategy and a prompt-edit co-iterative paradigm, coupled with a semantics-oriented taxonomy for motion descriptions—overcoming limitations of single-shot prompting. Keyframer integrates large language models’ (LLMs) natural language understanding and code generation capabilities with SVG specification parsing, real-time frontend editing, and variant synthesis mechanisms. A user study (n=13) demonstrates that Keyframer significantly improves design efficiency and creative ideation, empowering non-expert users to produce high-quality keyframe animations effectively.
This study addresses the issue that redundant views increase computational overhead and degrade reconstruction quality in feed-forward novel view synthesis. We propose a lightweight, rendering-free keyframe selector that integrates geometric and image features to construct a multi-criteria scoring system based on coverage, redundancy, and clarity. By employing a genetic algorithm for offline optimal subset search and knowledge distillation to train a compact network, our method breaks away from the conventional selection paradigm that relies on reconstruction feedback. Experiments across six datasets demonstrate that the proposed approach outperforms existing baselines while significantly reducing selection costs. Notably, the curated subsets achieve superior performance compared to full-sequence inputs and exhibit strong cross-paradigm generalization capabilities.
This study addresses the inherent trade-off between frame interaction timing and visual token cost in video encoding by proposing a novel paradigm that decouples the process into three stages: frame representation, token allocation, and temporal interaction. Methodologically, it adopts a strategy of preserving independent frame evidence before jointly allocating tokens and deferring temporal context integration, thereby preventing premature feature mixing. Architecturally, the proposed compact encoder combines a frozen image encoder, a question-aware selector, and a lightweight residual refinement module. Extensive experiments across thirteen benchmarks demonstrate that this approach matches full-image performance while utilizing only approximately 30% of the tokens, achieving substantial inference acceleration.
This study addresses the limitation of static keyframes in long-form video question answering, where they often fail to capture implicit information and handle complex queries. We propose a training-free dynamic visual clue retrieval framework that reformulates questions into dynamic clues. By leveraging vision-language models (VLMs) and agent-based reasoning, the method iteratively refines these clues and autonomously re-explores the video. Furthermore, it integrates dynamic prompt decomposition with budget allocation to achieve evidence-driven keyframe selection, thereby overcoming static context constraints. Experimental results demonstrate that our framework achieves state-of-the-art performance across 27 scenarios on three benchmarks, yielding an average improvement of 4.54%. Notably, this approach requires only two VLM invocations while attaining 92% of the maximal similarity gain, highlighting its efficiency and effectiveness for long-video understanding.
研究通过选择、压缩和再投资视觉标记的方法,优化长视频多模态语言模型中的帧选择问题,提高模型性能。
This study addresses the challenges of historical frame redundancy and poor future consistency caused by current-content-based selection in long video generation. To this end, we propose a prospective frame selector that introduces a novel future-aware historical selection mechanism. By predicting compact tokens representing future information demands, the method explicitly retrieves key historical frames to guide generation, employing a lightweight inference architecture for plug-and-play integration across diverse models. Extensive experiments conducted on five benchmarks with eleven models demonstrate that the proposed approach significantly enhances long-range temporal consistency, visual quality, and motion alignment.