Score
Design, implement, and evaluate temporal attention mechanisms — including temporal self-attention encoders, attention-based pooling layers, and time-aware attention modules — that compute attention weights over positions in a sequence to capture temporal dependencies and motion-related patterns. Integrate these components into classifiers and representation modules that produce sequence-level predictions and attention scores usable for interpretation, pooling, and downstream tasks.
This paper investigates whether Randomized Time Warping (RTW) can be formally characterized as a self-attention mechanism and systematically compares its modeling capacity and computational efficiency against Transformer self-attention for action recognition. Methodologically, we establish, for the first time, a theoretical equivalence between RTW and global, non-parametric self-attention; further, we propose an RTW-DTW joint framework that jointly performs nonlinear sequence alignment and discriminative feature enhancement. Our key contributions are threefold: (1) We reveal that RTW inherently enables implicit global temporal modeling—circumventing the locality bias imposed by Transformer’s quadratic-complexity constraints; (2) We empirically demonstrate strong correlation (mean Spearman’s ρ = 0.80) between RTW-derived attention weights and standard self-attention weights; (3) On Something-Something V2, RTW achieves a +5% accuracy gain over the baseline Transformer, while simultaneously attaining superior computational efficiency and performance.
Current video large language models (Video-LLMs) face significant bottlenecks in modeling complex temporal dynamics—such as action evolution and inter-frame dependencies—hindering fine-grained temporal reasoning. To address this, we propose a novel architecture that pioneers the integration of stacked temporal attention modules directly into the visual encoder, enabling explicit time-structure and action-sequence modeling at the early visual representation stage. This design synergizes with a multimodal fusion mechanism to enhance inter-frame relational modeling and cross-modal alignment. Evaluated on VITATECS, MVBench, and Video-MME benchmarks, our model achieves an average performance gain of +5.5%, with particularly strong improvements in action recognition and temporal question answering—outperforming state-of-the-art methods. Our core contribution lies in fundamentally shifting temporal modeling to the foundational layer of the visual encoder, thereby overcoming the limitations of conventional Video-LLMs that rely solely on post-hoc fusion or lightweight temporal modules.
Existing Mamba models exhibit limited capability in modeling nonlinear dependencies and suffer from restricted receptive fields due to convolutional operations in time-series modeling. To address these limitations, this paper proposes Attention Mamba—a novel framework integrating state-space models with attention mechanisms. It introduces an adaptive pooling attention mechanism that incorporates global contextual information while reducing computational complexity; additionally, it designs a bidirectional Mamba block enabling end-to-end mapping from input to value representations, thereby overcoming local receptive field constraints. By deeply fusing state-space dynamics with attention-based global interaction, the framework significantly enhances long-range dependency capture and nonlinear modeling capacity. Extensive experiments across diverse time-series forecasting benchmarks demonstrate that Attention Mamba consistently outperforms Transformer, Informer, and standard Mamba variants, validating its effectiveness, robustness, and generalizability.
Standard scaled dot-product attention lacks explicit modeling of continuous monotonic alignment, limiting performance on frame-synchronous tasks such as text-to-speech (TTS). To address this, we propose stochastic clock attention: it models source–target sequence alignment as the meeting probability of two learned non-negative stochastic clocks, and derives a closed-form Gaussian scoring function via path integral theory—ensuring causality, smoothness, and near-diagonal preference. This mechanism intrinsically enforces continuous monotonic alignment without positional regularization, supports both normalized and unnormalized forms, and unifies parallel and autoregressive decoding. In TTS, it significantly improves alignment stability and robustness to global temporal scaling variations, while maintaining or surpassing baseline models in speech quality.
This work addresses the challenge of enhancing neural networks’ ability to focus on salient information in long-sequence and multimodal tasks. By establishing a unified theoretical framework for attention mechanisms, the study systematically analyzes their mathematical foundations, computational properties, and cross-task generalizability. The framework is instantiated across diverse architectures—including autoregressive Transformers, bidirectional encoders, Vision Transformers, and cross-modal attention models—demonstrating consistent performance gains. The research further uncovers an intrinsic relationship between attention structure and model interpretability, validates empirical scaling laws governing training dynamics and performance, and achieves state-of-the-art results on multiple benchmark datasets. Attention visualization techniques are employed to enhance model transparency, offering insights into the decision-making process of these architectures.
This study addresses the unclear formation mechanisms of temporal attention in video diffusion models, where population-level averaging obscures sparse specialization phenomena. Through checkpoint-level analysis, we systematically survey temporal attention heads during Open-Sora training. We introduce an entropy-normalized cross-frame attention concentration metric alongside preregistered selection rules, employing change-point detection, ablation studies, and correlation analyses. Results reveal that approximately 4–13% of attention heads specialize into local frame-routing patterns as early as the first temporal block, converging toward fixed structures and exposing sparse temporal organization principles masked by averaging. Although a causal link to generation quality remains unestablished, this work provides a transferable analytical paradigm for understanding the internal mechanisms of video diffusion models.
研究通过引入TimeCatch方法检测视觉-语言模型在视频序列中对时间一致性的敏感度,揭示了模型在帧级异常检测上的成功与时间异常检测上的不足。
Existing video understanding methods struggle to effectively model complex temporal dynamics. To address this limitation, this work presents the first systematic exploration of high-order spatiotemporal self-similarity (STSS) and introduces a lightweight, general-purpose Multi-Order Self-Similarity (MOSS) module. MOSS enhances action modeling by learning and fusing STSS features across multiple orders. Implemented with neural networks for efficient feature extraction and integration, the proposed module consistently achieves significant performance gains across diverse benchmarks, including action recognition, motion-oriented video question answering, and real-world robotic tasks, thereby demonstrating its broad applicability and effectiveness.
Current video models exhibit limitations in temporal understanding and heavily rely on large-scale datasets with language supervision, resulting in high training costs and constrained concept learning. This work proposes motion as a core modality, introducing point trajectories—structured motion cues—as an independent input for the first time. By employing a masked autoencoder to reconstruct occluded trajectories, the method enables self-supervised video representation learning without requiring language annotations or extensive appearance-based data. This approach substantially enhances temporal perception and data efficiency. The learned TIME embeddings achieve state-of-the-art performance in zero-shot settings, using four orders of magnitude less training data than existing methods.
This work addresses the ambiguity regarding whether existing single-stage video object detectors genuinely leverage temporal context, as standard evaluation metrics often fail to reveal their actual reliance on temporal information. To this end, we propose TemporalLens, a diagnostic framework that quantifies a model’s temporal dependency through controlled perturbations—including temporal shuffling, structured occlusion, and redundancy injection. Furthermore, we design YOLO-3D based on YOLOv8, explicitly preserving the temporal dimension within the backbone to enhance genuine temporal reasoning. Experiments demonstrate that TemporalLens effectively distinguishes between stacked 2D models and true temporal architectures, while YOLO-3D achieves an average mAP@50 improvement of 3.7 percentage points with 32-frame inputs, underscoring the critical role of temporal depth in performance gains.