๐ค AI Summary
This work addresses the high computational cost and reliance on offline tracking in existing depth-aware video panoptic segmentation methods, which hinder real-time deployment in autonomous driving. To overcome these limitations, the authors propose a unified online framework that integrates semantic, instance, and metric depth estimation within a single forward pass. The approach employs an Explicit Scene Discretization (ESD) mechanism to distinguish foreground from background and introduces a Discrete-to-Continuous (D2C) depth decoder to recover accurate metric depth. Furthermore, an Online Majority Voting (OMV) mechanism is incorporated to enhance temporal consistency and classification robustness. This method achieves state-of-the-art performance on Cityscapes-DVPS and SemKITTI-DVPS while significantly reducing latency, marking the first efficient solution for online 4D scene understanding suitable for real-time autonomous perception.
๐ Abstract
Safe autonomous navigation requires a holistic understanding of dynamic environments, necessitating the simultaneous estimation of metric depth, semantic segmentation, and instance trajectories. While depth-aware video panoptic segmentation (DVPS) unifies these tasks, existing approaches often rely on computationally expensive, multi-stage pipelines or offline tracking, rendering them unsuitable for real-time decision-making. To address this, we propose DVPSFormer, a unified online architecture designed for efficient 4D scene understanding. Central to our approach is explicit scene discretization (ESD), a novel mechanism that leverages segmentation queries to represent foreground and background regions, enabling a discrete-to-continuous (D2C) depth head to decode metric depth in a single pass. This tightly couples semantic and geometric learning while significantly reducing latency. Furthermore, we propose an online majority voting (OMV) mechanism that exploits temporal consistency to refine classification during instance tracking. DVPSFormer establishes a new state-of-the-art on the Cityscapes-DVPS and SemKITTI-DVPS benchmarks, offering a streamlined solution for online robotic perception. Code and models are available at https://royyang0714.github.io/DVPSFormer.