DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving

๐Ÿ“… 2026-07-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the high computational cost and reliance on offline tracking in existing depth-aware video panoptic segmentation methods, which hinder real-time deployment in autonomous driving. To overcome these limitations, the authors propose a unified online framework that integrates semantic, instance, and metric depth estimation within a single forward pass. The approach employs an Explicit Scene Discretization (ESD) mechanism to distinguish foreground from background and introduces a Discrete-to-Continuous (D2C) depth decoder to recover accurate metric depth. Furthermore, an Online Majority Voting (OMV) mechanism is incorporated to enhance temporal consistency and classification robustness. This method achieves state-of-the-art performance on Cityscapes-DVPS and SemKITTI-DVPS while significantly reducing latency, marking the first efficient solution for online 4D scene understanding suitable for real-time autonomous perception.
๐Ÿ“ Abstract
Safe autonomous navigation requires a holistic understanding of dynamic environments, necessitating the simultaneous estimation of metric depth, semantic segmentation, and instance trajectories. While depth-aware video panoptic segmentation (DVPS) unifies these tasks, existing approaches often rely on computationally expensive, multi-stage pipelines or offline tracking, rendering them unsuitable for real-time decision-making. To address this, we propose DVPSFormer, a unified online architecture designed for efficient 4D scene understanding. Central to our approach is explicit scene discretization (ESD), a novel mechanism that leverages segmentation queries to represent foreground and background regions, enabling a discrete-to-continuous (D2C) depth head to decode metric depth in a single pass. This tightly couples semantic and geometric learning while significantly reducing latency. Furthermore, we propose an online majority voting (OMV) mechanism that exploits temporal consistency to refine classification during instance tracking. DVPSFormer establishes a new state-of-the-art on the Cityscapes-DVPS and SemKITTI-DVPS benchmarks, offering a streamlined solution for online robotic perception. Code and models are available at https://royyang0714.github.io/DVPSFormer.
Problem

Research questions and friction points this paper is trying to address.

depth-aware video panoptic segmentation
autonomous driving
online perception
real-time decision-making
4D scene understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Explicit Scene Discretization
Discrete-to-Continuous Depth
Online Majority Voting
Unified Online Architecture
Depth-aware Video Panoptic Segmentation
๐Ÿ”Ž Similar Papers
No similar papers found.