Structure over Depth: A Single-Block Spatio-Temporal Transformer for Multi-Entity Reasoning

📅 2026-07-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Modeling multivariate temporal data involving multiple entities poses significant challenges in efficiently capturing complex dependencies across entities, over time, and their intricate couplings. This work proposes a structure-aware single-block Transformer architecture that explicitly decomposes and models spatial, temporal, and cross-domain interactions through parallel spatial and temporal self-attention mechanisms coupled with bidirectional cross-attention. A learnable gating mechanism is introduced to effectively fuse information from multiple sources. By adopting explicit factorization instead of deep stacking, the proposed model achieves competitive or superior performance compared to more complex architectures across multiple tasks, despite having only 1.76 million parameters, thereby demonstrating its efficiency and effectiveness.
📝 Abstract
Modeling multi-entity temporal data requires capturing dependencies across entities, time, and their interactions. Transformer-based approaches perform well but often rely on deep stacks of layers to learn these heterogeneous dependencies implicitly, increasing computational cost. We revisit this problem from a structural perspective and decompose multi-entity temporal dynamics into three interaction types: spatial interactions among entities, temporal interactions across time, and cross interactions coupling the two domains. We propose a structured spatio-temporal transformer block that explicitly models all three within a single stage. It uses parallel spatial and temporal self-attention, followed by bidirectional cross-attention, and combines the outputs through learnable gated fusion. By directly encoding these complementary views, the model reduces the need for deep stacking. We evaluate the approach on video-based group activity recognition, skeleton-based human interaction analysis, and wearable sensor-based activity recognition. Despite its simplicity, the single structured Transformer block matches or outperforms deeper architectures with only 1.76M parameters. The results suggest that depth in prior models partly compensates for implicit and entangled interaction modeling, whereas explicit factorization offers a more efficient and transparent alternative. More broadly, this work supports a structure-first design principle: expressive multi-entity temporal reasoning can emerge by exposing interaction structure rather than relying on depth.
Problem

Research questions and friction points this paper is trying to address.

multi-entity temporal reasoning
spatio-temporal modeling
interaction decomposition
transformer architecture
computational efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

structured transformer
spatio-temporal modeling
multi-entity reasoning
explicit interaction factorization
gated fusion
N
Narthana Sivalingam
Department of Electrical and Electronic Engineering, University of Peradeniya, Peradeniya, Sri Lanka
S
Santhirarajah Sivasthigan
Department of Electrical and Electronic Engineering, University of Peradeniya, Peradeniya, Sri Lanka
Buddhi Wijenayake
Buddhi Wijenayake
Student at University of Peradeniya
Computer VisionStatistics
R
Roshan Godaliyadda
Department of Electrical and Electronic Engineering, University of Peradeniya, Peradeniya, Sri Lanka
Vijitha Herath
Vijitha Herath
Professor of Electrical and Electronic Engineering, University of Peradeniya, Sri Lanka
Remote SensingSpectral ImagingTelecommunicationsAI
P
Parakrama Ekanayake
Department of Electrical and Electronic Engineering, University of Peradeniya, Peradeniya, Sri Lanka