🤖 AI Summary
This work addresses the performance degradation in bird's-eye-view (BEV) semantic segmentation caused by intrinsic noise introduced during feature learning. It pioneers the integration of noise estimation principles from denoising diffusion models into BEV perception, proposing a general-purpose BEV feature denoising framework. The method explicitly models and removes noise in BEV features through a UNet-based noise estimation module, coupled with a task-decomposed training paradigm to enable effective supervision. Extensive experiments on the large-scale real-world nuScenes dataset demonstrate significant improvements in semantic segmentation accuracy across four state-of-the-art BEV models under three major view transformation paradigms, while also revealing three key design insights critical to effective denoising in BEV representation learning.
📝 Abstract
In this paper, we present a framework dubbed \textbf{BEV-Denoise} that estimates and removes intrinsic noise from learned Bird's-Eye-View (BEV) features to achieve accurate BEV semantic segmentation. Inspired by the noise estimation capability of Denoising Diffusion Probabilistic Models (DDPM), we design a UNet-based noise estimation module that learns to estimate the noise from the learned BEV features. The estimated noise is then subtracted from the BEV features and fed to BEV map decoders for the final prediction results. To facilitate supervision for the noise estimation module, we follow a sequential learning paradigm called Task Decomposition (TD) where a pre-trained BEV map autoencoder is employed to train a view transformation (VT) encoder. We share three key insights learned from our intensive experiments that are critical for improved performance. We apply our framework to four existing models, encompassing the three major VT paradigms. Experimental results on a large-scale real-world dataset, nuScenes, demonstrate the effectiveness of our framework.