Score
Designs and implements embedding representations and injection mechanisms that encode temporal position or object-track identity into video transformer architectures; builds and compares parametrizations (learned or fixed temporal/track embeddings) and integration strategies (addition, concatenation, attention bias, conditioning tokens) and analyzes their effects on temporal modeling, object persistence, and downstream video prediction/recognition tasks.
Current video models exhibit limitations in temporal understanding and heavily rely on large-scale datasets with language supervision, resulting in high training costs and constrained concept learning. This work proposes motion as a core modality, introducing point trajectories—structured motion cues—as an independent input for the first time. By employing a masked autoencoder to reconstruct occluded trajectories, the method enables self-supervised video representation learning without requiring language annotations or extensive appearance-based data. This approach substantially enhances temporal perception and data efficiency. The learned TIME embeddings achieve state-of-the-art performance in zero-shot settings, using four orders of magnitude less training data than existing methods.
Existing video classification methods naively average frame-level embeddings from pretrained Transformer encoders, neglecting critical temporal structures—including event ordering, dynamic feature importance, and duration variability. While mainstream temporal modeling approaches require architectural modifications and full retraining, they are incompatible with already fine-tuned large models. This paper proposes an encoder-agnostic, lightweight temporal matching framework: it maps fixed-length embeddings into variable-length multivariate time series and introduces a learnable per-frame, per-feature weighting mechanism. Inspired by time-series alignment, the framework employs a dedicated neural architecture for temporal modeling. It adds fewer than 1.8% parameters and requires no encoder modification or retraining. Evaluated on Something-Something V2, Kinetics-400, and HMDB51, our method achieves 77.2%, 89.1%, and 88.6% Top-1 accuracy, respectively, with training completed in under three hours.
Prior work on temporal video grounding lacks systematic evaluation of video encoders, leading to architecture-specific overfitting. Method: This paper conducts the first end-to-end comparative study of three dominant video encoder paradigms—CNNs, temporal reasoning models (e.g., RNNs), and Transformers—across three standard benchmarks (Charades-STA, ActivityNet-Captions, YouCookII), integrating each uniformly into canonical grounding architectures. Contribution/Results: We reveal complementary feature representations across encoder types and identify consistent, model-specific error patterns in localization. Empirical results demonstrate that encoder choice significantly impacts grounding accuracy, with distinct error distributions tied to architectural inductive biases. Our findings provide empirical grounding for video representation design and propose an encoder diversity principle to mitigate architectural overfitting—thereby enhancing model robustness and generalization in temporal video grounding.
Existing video understanding research overlooks how structural characteristics of datasets—such as motion complexity, temporal span, hierarchical composition, and multimodal richness—guide the evolution of model architectures. Method: We propose a dataset-centric analytical framework that systematically interprets mainstream architectures—including two-stream networks, 3D CNNs, RNNs, Transformers, and multimodal foundation models—as responses to dataset-imposed inductive biases. Our approach integrates literature review with architecture–bias–task alignment analysis, unifying inductive bias theory and multimodal learning paradigms. Contribution/Results: We establish, for the first time, a unified “dataset → inductive bias → model design” framework, revealing the intrinsic logic underlying architectural evolution. The framework yields principled, generalizable design guidelines for video understanding models that balance scalability and task adaptability, advancing both theoretical understanding and practical model development.
Existing custom video generation methods struggle to simultaneously preserve subject appearance fidelity and ensure temporal motion consistency, primarily due to the lack of object-level subject-motion disentangled modeling. This paper proposes a subject-motion representation disentanglement and alignment framework for text-to-video generation. We introduce the first object-level representation alignment mechanism, design a sparse spatiotemporal LoRA injection strategy to minimize fine-tuning interference, and develop a collaborative self-supervised subject encoder and optical-flow-based motion encoder. The method integrates self-supervised representation learning, optical-flow-driven motion modeling, efficient LoRA-based fine-tuning, and spatiotemporally sparse adapters. Evaluated on multiple benchmarks, it achieves significant improvements in subject similarity (+12.6%) and motion consistency (+9.8%), enabling fine-grained, disentangled controllable generation with both high visual fidelity and temporal stability.
This work addresses the limitations of existing video understanding methods in modeling long-range temporal dependencies. To overcome this challenge, the authors propose a systematic solution that integrates state-space layers and recurrent adapters to efficiently capture long-term temporal dynamics. They further introduce a fine-grained action-moment contrastive learning mechanism combined with a noise-robust training strategy to enhance the temporal reasoning capabilities of large vision-language models. The contributions include the release of two new long-form video benchmark datasets, empirical insights into the critical role of vision-language interfaces in temporal understanding, and significant improvements in modeling dynamic video content through parameter-efficient fine-tuning and temporally oriented training objectives.
This work proposes a novel self-supervised visual representation learning paradigm, Temporal Difference in Vision (TDV), which eschews strong inductive biases such as data augmentation, masking, or cropping commonly used in existing methods. Instead, TDV leverages the temporal causal assumption that “the past causes the future” in videos, jointly training an image encoder and a motion encoder so that the sum of the current frame’s representation and the motion representation approximates the representation of the subsequent frame. Relying solely on this weak temporal assumption, the method achieves state-of-the-art performance among self-supervised approaches on dense spatial tasks, demonstrating the effectiveness and potential of modeling temporal causality for large-scale visual representation learning.
This study systematically investigates video-text cross-modal representation alignment, focusing on modern encoders’ spatiotemporal modeling capabilities and their relationship with downstream performance. Method: We propose parameterized test-time scaling laws to quantitatively link semantic alignment degree with video understanding ability; design a novel temporal reasoning benchmark to overcome limitations of conventional zero-shot classification evaluation; and integrate multi-frame video encoding, text-set alignment, and regression-based modeling to jointly learn static and dynamic representations. Results: Experiments demonstrate that strong text alignment significantly enhances general-purpose video representation quality, and alignment metrics reliably predict model performance across diverse video understanding tasks. Our work advances the understanding of multimodal model internals and provides an interpretable pathway for alignment optimization.
This study investigates the feasibility of video content understanding using only camera motion trajectories—bypassing pixel-level processing. To this end, we propose CamFormer, a model that encodes sequences of camera poses into a joint embedding space aligned with natural language via contrastive learning, enabling cross-modal semantic alignment. Our key contribution is the first systematic validation that camera trajectories intrinsically encode rich semantic information: *how* a camera moves reliably reflects *what* action is occurring or *what* scene is being observed—establishing trajectory as a lightweight, robust, and general-purpose modality for video understanding. CamFormer is agnostic to pose estimation methods and achieves state-of-the-art performance on downstream tasks including cross-modal retrieval, action classification, and temporal reasoning. It demonstrates strong generalization across domains and robustness to modality variations, underscoring the viability of trajectory-centric video analysis.