SV-TAD: Native Sparse Convs for Efficient Temporal Action Detection

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in long video understanding where token pruning disrupts spatial structures, leading to computational redundancy and poor scalability in adapters. To overcome this, we propose SV-TAD, a framework that introduces native sparse 2D convolution primitives to break the reliance of conventional convolutions on complete spatial grids. This innovation enables lightweight adapters to operate efficiently and directly on dynamically pruned token sets without dense reconstruction, while naturally accommodating auxiliary task tokens. Built upon backbones such as VideoMAEv2-L, the proposed method maintains state-of-the-art accuracy on benchmarks including THUMOS-14, while reducing computation by 64% and achieving a 2.2× speedup. Ultimately, SV-TAD realizes simultaneous optimization of both parameter efficiency and inference speed for temporal action detection.
📝 Abstract
To adapt billion-parameter Vision Transformers for long-video understanding, recent methods freeze the backbone and train lightweight convolutional modules. While effective for parameter-efficient training, existing adapters do not reduce inference-time computation, leaving scalability with respect to video length largely unaddressed. Token selection can reduce attention cost by pruning redundant tokens, but it breaks the spatial grid structure required by convolutional adapters. This forces an expensive dense reconstruction, nullifying much of the potential speedup. We address this by introducing native sparse 2D convolutions, a primitive that allows these adapters, for the first time, to operate directly and efficiently on dynamically pruned token sets. We integrate this primitive into SV-TAD, an adapter framework for temporal action detection, reducing VideoMAEv2-L computation by up to 64% and achieving 2.2x faster inference, while maintaining state-of-the-art accuracy on THUMOS-14 and ActivityNet-1.3. When scaled to InternVideoNext-L, our approach surpasses the previous state of the art at roughly half its computational cost. Moreover, the sparse formulation naturally supports auxiliary task tokens, which improves fine-grained assembly detection on ATTACH.
Problem

Research questions and friction points this paper is trying to address.

Temporal Action Detection
Sparse Convolutions
Token Pruning
Parameter-Efficient Adaptation
Long-Video Understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Native Sparse Convolutions
Temporal Action Detection
Parameter-Efficient Adapter
Token Pruning
Long-Video Understanding
🔎 Similar Papers
No similar papers found.
R
Ricardo Pizarro
Universidad de Alcalá, Alcalá de Henares, Spain
R
Roberto Valle
Universidad Politécnica de Madrid, Madrid, Spain
J
José M. Buenaposada
Universidad Rey Juan Carlos, Móstoles, Spain
L
Luis M. Bergasa
Universidad de Alcalá, Alcalá de Henares, Spain
Luis Baumela
Luis Baumela
Departamento de Inteligencia Artificial, Universidad Politécnica de Madrid
computer vision