Attention Mechanism in Randomized Time Warping

📅 2025-08-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper investigates whether Randomized Time Warping (RTW) can be formally characterized as a self-attention mechanism and systematically compares its modeling capacity and computational efficiency against Transformer self-attention for action recognition. Methodologically, we establish, for the first time, a theoretical equivalence between RTW and global, non-parametric self-attention; further, we propose an RTW-DTW joint framework that jointly performs nonlinear sequence alignment and discriminative feature enhancement. Our key contributions are threefold: (1) We reveal that RTW inherently enables implicit global temporal modeling—circumventing the locality bias imposed by Transformer’s quadratic-complexity constraints; (2) We empirically demonstrate strong correlation (mean Spearman’s ρ = 0.80) between RTW-derived attention weights and standard self-attention weights; (3) On Something-Something V2, RTW achieves a +5% accuracy gain over the baseline Transformer, while simultaneously attaining superior computational efficiency and performance.

Technology Category

Computer Vision: Motion & TrackingKnowledge Representation and Reasoning: Geometric, Spatial, and Temporal ReasoningPlanning, Routing, and Scheduling: Activity and Plan Recognition

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsGraph Algorithms and Modeling for the Web: Algorithms and analysis for heterogeneous, signed, attributed, multi-relational, temporal, higher-order, and annotated Web-related graphs
📝 Abstract
This paper reveals that we can interpret the fundamental function of Randomized Time Warping (RTW) as a type of self-attention mechanism, a core technology of Transformers in motion recognition. The self-attention is a mechanism that enables models to identify and weigh the importance of different parts of an input sequential pattern. On the other hand, RTW is a general extension of Dynamic Time Warping (DTW), a technique commonly used for matching and comparing sequential patterns. In essence, RTW searches for optimal contribution weights for each element of the input sequential patterns to produce discriminative features. Although the two approaches look different, these contribution weights can be interpreted as self-attention weights. In fact, the two weight patterns look similar, producing a high average correlation of 0.80 across the ten smallest canonical angles. However, they work in different ways: RTW attention operates on an entire input sequential pattern, while self-attention focuses on only a local view which is a subset of the input sequential pattern because of the computational costs of the self-attention matrix. This targeting difference leads to an advantage of RTW against Transformer, as demonstrated by the 5% performance improvement on the Something-Something V2 dataset.
Problem

Research questions and friction points this paper is trying to address.

RTW interprets sequential pattern weights as self-attention
Compares RTW and self-attention mechanisms in motion recognition
Demonstrates RTW's performance advantage over Transformer models
Innovation

Methods, ideas, or system contributions that make the work stand out.

RTW interpreted as self-attention mechanism
RTW searches optimal weights for sequential elements
RTW operates on entire input pattern unlike local self-attention
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yutaro Hiraoka
The Japan Research Institute, Limited
Kazuya Okamura
Kazuya Okamura
Graduate School of Science and Technology, University of Tsukuba
K
Kota Suto
Graduate School of Science and Technology, University of Tsukuba
K
Kazuhiro Fukui
Tsukuba Institute for Advanced Research, Department of Computer Science, University of Tsukuba