Pocket-STVG: lightweight architecture for Spatio-Temporal Video Grounding

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the heavy reliance on large models and high computational costs in spatiotemporal video grounding by proposing a lightweight cascaded architecture. Methodologically, it integrates MobileViCLIP for temporal encoding, MDETR for spatial encoding and decoding, and a 1D U-Net, while introducing a decoupled precomputation indexing pipeline that uniformly supports both weakly supervised and zero-shot scenarios. With fewer than 90 million parameters, the proposed architecture achieves performance comparable to existing weakly supervised methods while substantially reducing GPU memory consumption and computational overhead. This work establishes a new paradigm for efficient spatiotemporal tube grounding in videos.
📝 Abstract
Spatio-Temporal Video Grounding (STVG) aims to localize the spatio-temporal tube in a video corresponding to a natural language query. While recent methods achieve strong performance in fully supervised, weakly supervised, and zero-shot settings, they typically rely on computationally expensive architectures, complex training pipelines, or multimodal large language models. We present Pocket-STVG (P-STVG), a lightweight cascade architecture that addresses STVG by combining efficient pre-trained components instead of large end-to-end models. P-STVG integrates a temporal-aware video encoder based on MobileViCLIP, a spatial encoder-decoder derived from MDETR, and a shared aligned text encoder. Temporal localization is performed through either a lightweight 1D U-Net or a simple thresholding strategy, enabling the same framework to operate in both weakly supervised and zero-shot settings. Furthermore, video representations are precomputed independently of the query, yielding an indexing-friendly pipeline for efficient inference and large-scale video collections. Despite requiring fewer than 90M parameters, P-STVG performs on par with weakly supervised methods and improves on earlier zero-shot approaches at a fraction of their memory and computational cost, establishing a favorable performance-efficiency trade-off for STVG.
Problem

Research questions and friction points this paper is trying to address.

Spatio-Temporal Video Grounding
computational efficiency
lightweight architecture
zero-shot
weakly supervised
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatio-Temporal Video Grounding
Lightweight Architecture
Zero-shot Learning
Efficient Inference
Cascade Architecture
🔎 Similar Papers
No similar papers found.
Alberto Presta
Alberto Presta
PhD Student at university of Turin, computer science department
Computer visionArtificial intelligence
Michal Byra
Michal Byra
Polish Academy of Sciences, Samsung AI Center, RIKEN
biomedical image analysisimaging sciencesneural networksrepresentation learning
G
Grzegorz Stefański
Samsung AI Center, Warsaw, Poland
K
Karol Szurkowski
Samsung AI Center, Warsaw, Poland
E
Eryk Kołodziejczyk
Samsung AI Center, Warsaw, Poland
K
Krzysztof Arendt
Samsung AI Center, Warsaw, Poland