quasi-gaussian sampling

Design and implement frame‑sampling algorithms that assume a quasi‑Gaussian distribution of frame scores to select representative video frames within an estimated 3‑σ interval, adapting the sampling interval per local or global query. Build training‑free, single‑hyperparameter methods that pick frames under a fixed budget to reduce computation and memory footprint.

quasi-gaussiansampling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Scalable Frame Sampling for Video Classification: A Semi-Optimal Policy Approach with Reduced Search Space

Sep 09, 2024
JL
Junho Lee
🏛️ Seoul National University | LG AI Research | Twelve Labs | Google Research

To address the computationally intractable explosion of the frame-sampling search space—combinatorially scaling as $inom{T}{N}$—in video classification, this work proposes a decoupled single-frame value assessment framework. It reformulates frame selection as an independent confidence scoring problem and approximates the optimal $N$-frame subset via greedy selection of the top-$N$ highest-scoring frames. This reduces time complexity from $O(T^N)$ to $O(T)$, and for the first time provides theoretical guarantees on approximation optimality and scalability. The method is lightweight, model-agnostic, and requires no fine-tuning of downstream classifiers. Extensive experiments across multiple benchmarks and architectures demonstrate consistent superiority over state-of-the-art sampling methods, robustness to variations in both frame count $N$ and video length $T$, and over 100× inference speedup.

Approximating optimal policy with reduced complexityMaximizing video classifier performance efficientlyReducing vast search space in video frame sampling

This work addresses frame redundancy in long videos by proposing a lightweight, training-free frame sampling method inspired by the brain’s predictive coding mechanism, requiring neither auxiliary networks nor video-specific hyperparameter tuning. The approach models a video as a differentiable trajectory in a visual latent space and employs Taylor expansion to predict the evolution of frames along this trajectory. Sampling is achieved by identifying “temporal surprises”—frames that significantly deviate from the predicted path—based on analyses of velocity and acceleration in the feature trajectory. With computational overhead merely 0.02× that of baseline methods (30× lower), the technique substantially outperforms query-agnostic strategies such as uniform sampling across three long-form video question-answering benchmarks and ten downstream tasks, achieving up to a 12.5 percentage point accuracy gain under constrained frame budgets.

information-rich momentslong-form videopredictive deviation

MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMs

Jan 06, 2025
HS
Hui Sun
🏛️ Nanjing University | Alibaba

To address the performance degradation of Video-LLMs caused by suboptimal keyframe selection in long videos, this paper proposes MDP3—a training-free, model-agnostic keyframe selection method. MDP3 jointly optimizes for query relevance, list-level diversity, and temporal coherence under a unified framework. It is the first approach to integrate: (i) conditional Gaussian kernels for query-aware frame similarity estimation; (ii) determinantal point processes (DPPs) to model inter-frame diversity; and (iii) a Markov decision process coupled with dynamic programming to enforce temporal contiguity. Theoretical analysis provides a $(1-1/e)$-approximation guarantee for the resulting combinatorial optimization. Extensive experiments on multiple video understanding benchmarks demonstrate that MDP3 significantly outperforms uniform sampling, query-matching heuristics, and other baselines—achieving substantial accuracy gains while maintaining high computational efficiency and strong cross-model generalizability.

Diverse Frame SelectionSequential ConsistencyVideo Information Retrieval

This study addresses the inefficiency of frame sampling in long video question answering caused by fixed global budgets. To overcome this limitation, we propose a training-free, plug-and-play dynamic frame selection strategy. Specifically, this work pioneers the integration of a dynamic budget adjustment mechanism atop existing selectors to transcend fixed constraints, coupled with a training-free meta-sampling algorithm that optimizes the frame sequences fed into multimodal large language models. Experimental results demonstrate that the proposed strategy reduces the number of input frames by an average of 8.9% while simultaneously improving accuracy across most configurations, thereby significantly enhancing computational efficiency.

EfficiencyFrame SelectionLong-Video Question Answering

Existing single-frame-to-video diffusion models for two-frame-constrained video keyframe interpolation suffer from out-of-manifold generation and artifacts due to parallel bidirectional sampling, while relying on multiple denoising iterations. Method: We propose a sequential bidirectional diffusion sampling strategy—enabling the first single-pass forward–backward cascaded sampling—to eliminate redundant denoising. Our approach integrates CFG++ classifier-free guidance with DDS (Dynamic Denoising Scheduling) to enhance temporal coherence and in-manifold generation fidelity. Contribution/Results: On a single RTX 3090 GPU, our method generates high-quality 25-frame interpolated videos at 1024×576 resolution in just 195 seconds, achieving state-of-the-art performance in keyframe interpolation while significantly improving efficiency and visual quality.

Address off-manifold issues in two-frame conditioned video generation.Enhance video interpolation using bidirectional diffusion sampling.Improve quality and efficiency of intermediate frame generation.

Latest Papers

What's happening recently
View more

This work addresses the high computational and memory costs that hinder multimodal large language models in long video understanding, as well as the inflexibility and noise sensitivity of existing keyframe sampling methods. The authors formalize video frame sampling as a quasi-Gaussian distribution problem and propose AdaQ, a training-free adaptive sampling method that dynamically adjusts the sampling interval based on the 3σ rule. AdaQ balances local and global query requirements using only a single hyperparameter. Integrated with multi-visual embeddings and the Qwen3-VL-8B model, AdaQ achieves an average performance gain of 15.8% over GPT-4o while sampling merely 64 frames, significantly outperforming current keyframe selection approaches.

Adaptive samplingComputational efficiencyKeyframe selection

This work addresses the high computational cost of video large language models caused by processing dense frames, as well as the limitations of existing keyframe selection methods that often fall into local optima and introduce noisy frames. To this end, the authors propose a training-free frame selection framework that introduces, for the first time, a “directed diversity” metric to unify relevance and diversity into an irreplaceability score. Coupled with a budget-aware adaptive iterative mechanism, the method dynamically optimizes temporal context coverage while preserving core semantics. Evaluated on the LLaVA-Video-7B model across long-video benchmarks, the approach achieves an average performance gain of 12.5%, significantly outperforming baseline strategies such as uniform sampling.

Computational CostFrame SelectionKeyframe Sampling

Existing frame sampling methods for long video understanding struggle to balance global coverage with the capture of transient critical events, limiting downstream task performance. This work proposes InfoShot—a task-agnostic, shot-aware frame sampler that first segments videos into semantically coherent clips using shot boundary detection and then selects two complementary keyframes from each clip: one representing dominant content and the other capturing anomalous changes. By optimizing an information-theoretic objective, InfoShot preserves both structural context and sparse biases without requiring model retraining. To evaluate short-term anomaly detection, we introduce SynFlash, a controllable synthetic benchmark. Experiments demonstrate that under strict frame budgets, InfoShot significantly improves anomaly hit rates and video question-answering accuracy, achieving competitive or superior performance against strong baselines on standard benchmarks.

critical eventslong-video understandingshot-aware

This work addresses the challenges of inefficient reasoning and missed critical frames in long-form video question answering. The authors propose a question-adaptive greedy frame selection method that jointly optimizes query relevance and semantic representativeness under a fixed frame budget. Frame quality is evaluated in dual embedding spaces using SigLIP and DINOv2, while a facility location coverage term enhances diversity. The approach employs a normalized, monotonic, and submodular objective function, guaranteeing a (1−1/e)-approximation. Additionally, a lightweight text-only question-type classifier dynamically selects the optimal sampling strategy. Evaluated on the MLVU dataset, the method significantly outperforms uniform sampling and existing strong baselines, with particularly notable gains under low frame budgets.

frame selectionlong video understandingquery relevance

Hot Scholars

YL

Yanxiao Li

National Energy Technology Laboratory
DP

David Pujol

Tumult Labs
PrivacyAlgorithmic fairness
HL

Haochen Li

Tsinghua university
cell-cell communicationsingle-cell genomicsspatial transcriptomics