LVMT: Video Mask Transformer for Long-term Video Segmentation

๐Ÿ“… 2026-09-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the rigid temporal information filtering in online video segmentation under long-term occlusion, as well as memory overflow and vanishing gradients during long video training. To tackle these challenges, we propose the Long-term Video Mask Transformer (LVMT). Specifically, we introduce a lightweight GRU-based adaptive temporal propagation module to enable dynamic filtering of historical information. Furthermore, we present a Truncated Query Propagation (TQP) strategy that overcomes long-sequence training bottlenecks through chunked processing and local backpropagation. Extensive experiments demonstrate that LVMT establishes new state-of-the-art performance across six benchmarks while achieving a tenfold inference speedup compared to previous best models. These results indicate that our approach significantly enhances long-term tracking robustness in complex occlusion scenarios.
๐Ÿ“ Abstract
Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlusions. We hypothesize that this limitation is caused by (i) the inability of their temporal propagation mechanism to adaptively select the object information that is propagated across time, and (ii) their inability to be trained on long videos due to memory requirements and vanishing gradients. To address the first limitation, we propose to use a lightweight GRU-based temporal propagation module that can learn to select which information it keeps in memory and propagates across time. Second, to allow training on long videos, we introduce Truncated Query Propagation (TQP), a training strategy in which the model processes a video in chunks of frames, where information about tracked objects is propagated between chunks but backpropagation is only conducted in individual chunks, enabling longer temporal supervision without out-of-memory issues, inference overhead, or vanishing gradients. The resulting model is called the Long-term Video Mask Transformer (LVMT). Extensive experiments on six benchmarks show that LVMT sets a new state of the art across a range of video segmentation tasks, while retaining the speed of the highly efficient model it is based on, making it 10X faster than the prior state of the art. Code: https://www.tue-mps.org/lvmt
Problem

Research questions and friction points this paper is trying to address.

video segmentation
long-term occlusion
temporal propagation
object tracking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long-term Video Segmentation
Video Mask Transformer
GRU-based Temporal Propagation
Truncated Query Propagation
Online Video Segmentation
๐Ÿ”Ž Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30