bev feature extraction

Designs, builds, or analyzes methods that extract and represent bird’s‑eye‑view (BEV) feature maps from sensor-derived features, including projecting image or point features into BEV, encoding temporal context, and producing dense or sparse scene‑level BEV representations. Develops and evaluates denoising and refinement procedures that estimate and remove intrinsic noise, fuse probabilistic object primitives, and clean or project BEV features for use by downstream scene‑level perception modules.

bevfeatureextraction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.08
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the performance degradation in bird's-eye-view (BEV) semantic segmentation caused by intrinsic noise introduced during feature learning. It pioneers the integration of noise estimation principles from denoising diffusion models into BEV perception, proposing a general-purpose BEV feature denoising framework. The method explicitly models and removes noise in BEV features through a UNet-based noise estimation module, coupled with a task-decomposed training paradigm to enable effective supervision. Extensive experiments on the large-scale real-world nuScenes dataset demonstrate significant improvements in semantic segmentation accuracy across four state-of-the-art BEV models under three major view transformation paradigms, while also revealing three key design insights critical to effective denoising in BEV representation learning.

BEV FeaturesBird's-Eye-ViewIntrinsic Noise

OnlineBEV: Recurrent Temporal Fusion in Bird's Eye View Representations for Multi-Camera 3D Perception

Jul 11, 2025
JK
Junho Koh
🏛️ Hyundai Motors | Hanyang University | Seoul National University

In multi-view pure-vision 3D perception, dynamic objects impede temporal alignment of bird’s-eye-view (BEV) features across frames. To address this, we propose the Motion-guided BEV Fusion Network (MBFNet), which introduces a recurrent online temporal fusion architecture that explicitly models object motion and achieves spatiotemporal alignment of BEV features across frames. MBFNet incorporates a motion-guided dynamic alignment mechanism and an end-to-end differentiable temporal consistency loss to enhance fusion quality between historical and current BEV features. Crucially, the method requires no auxiliary motion annotations and is trained end-to-end solely from image inputs. Evaluated on the nuScenes benchmark, MBFNet achieves 63.9% NDS—the highest performance among pure-vision 3D detectors at the time—demonstrating for the first time the critical role of efficient online temporal fusion in learning robust dynamic BEV representations.

Align dynamic BEV features over time for accuracyEnhance 3D perception by fusing sequential BEV featuresImprove camera-only 3D object detection performance

Evaluating the Impact of Weather-Induced Sensor Occlusion on BEVFusion for 3D Object Detection

Nov 06, 2025
SK
Sanjay Kumar
🏛️ University of Limerick | Queen Mary University of London

This study systematically investigates the impact of weather-induced sensor occlusion on the performance of BEVFusion, a multimodal 3D object detector. Using the nuScenes dataset, we quantify the independent and fused performance of camera and LiDAR modalities in the bird’s-eye view (BEV) space under varying occlusion levels, evaluating with mean Average Precision (mAP) and NuScenes Detection Score (NDS). Results reveal strong LiDAR dependency: severe LiDAR occlusion degrades mAP by 47.3%, whereas moderate camera occlusion reduces mAP by 41.3%. In fusion mode, camera occlusion impact is markedly alleviated (mAP ↓4.1%), yet LiDAR occlusion still incurs a 26.8% mAP drop. This work provides the first empirical characterization of BEVFusion’s modality-specific vulnerability under adverse weather conditions, demonstrating that unimodal robustness is insufficient and underscoring the necessity of occlusion-resilient, adaptive fusion mechanisms. Our findings offer critical evidence for enhancing the reliability of autonomous driving perception systems in real-world, weather-impacted environments.

Analyzing sensor fusion robustness for autonomous vehicle perception systemsEvaluating weather-induced sensor occlusion effects on BEVFusion 3D detectionInvestigating camera and LiDAR performance degradation under occlusion conditions

OA-BEV: Bringing Object Awareness to Bird's-Eye-View Representation for Multi-Camera 3D Object Detection

Jan 13, 2023
XC
Xiaomeng Chu
🏛️ University of Science and Technology of China (USTC) | The University of Sydney (USYD)

To address feature distortion and boundary ambiguity in multi-camera BEV 3D detection—stemming from the ill-posed image-to-3D mapping—this paper proposes an object-aware pseudo-3D–depth joint modeling framework. Our method introduces two key innovations: (1) a novel object-level depth supervision scheme coupled with pseudo-voxel encoding, enhancing structural consistency in depth estimation; and (2) a 2D detection-guided foreground pixel projection mechanism integrated with deformable attention fusion, enabling precise spatial alignment and contextual enhancement. These components collectively improve object structural representation and spatial localization accuracy in the BEV feature space. Evaluated on nuScenes, our approach achieves 72.5% mAP and 78.3% NDS, outperforming state-of-the-art baselines including BEVDet and BEVFormer.

Addresses feature clutter in multi-camera 3D object detectionEnhances 3D detection accuracy using object-aware depth featuresImproves object-background differentiation in BEV representations

Benchmarking and Improving Bird's Eye View Perception Robustness in Autonomous Driving

May 27, 2024
SX
Shaoyuan Xie
🏛️ University of California, Irvine | National University of Singapore | Shanghai AI Laboratory | Nanyang Technological University

To address the insufficient robustness of bird’s-eye-view (BEV) perception in autonomous driving under atypical and challenging conditions—such as camera degradation and multi-sensor failures—this paper introduces RoboBEV, the first out-of-distribution (OOD) evaluation benchmark for BEV perception covering four tasks: 3D detection, semantic segmentation, depth estimation, and occupancy prediction, with systematic evaluation of 33 models. We propose a multi-granularity BEV robustness evaluation framework, revealing for the first time the weak correlation between standard performance and robustness. Additionally, we design a CLIP-based temporal enhancement strategy and a depth-agnostic BEV transformation method. Experiments demonstrate that pretraining and depth-agnostic BEV representations significantly improve robustness; temporal fusion boosts average robustness by 12.6%; and our CLIP-based strategy reduces mAP degradation by 37% under severe sensor degradation.

Autonomous VehiclesBird's Eye ViewStability and Performance Enhancement

Latest Papers

What's happening recently
View more

This work addresses the challenges of multi-camera data fusion and insufficient bird’s-eye-view (BEV) segmentation accuracy in complex driving scenarios by proposing a Transformer-based Variational Bifurcated network (TVB). TVB is the first approach to integrate variational inference with normalizing flows for BEV segmentation. It implicitly learns the mapping from multi-view images to a unified BEV representation through posterior BEV supervision, generating multiple candidate maps. A novel BEV-attention fusion module adaptively aggregates these candidates to enhance the realism and expressiveness of the resulting map. Evaluated on the nuScenes and OPV2V datasets, the proposed method significantly outperforms existing approaches in both multi-camera BEV segmentation and lane perception tasks, achieving state-of-the-art performance.

Autonomous DrivingBird's Eye View SegmentationEnvironmental Perception

This work addresses the performance degradation of existing BEV-based 3D object detection methods—primarily designed for pinhole cameras—when applied to fisheye cameras due to severe radial distortion, and the absence of benchmarks or effective solutions for hybrid camera setups. To bridge this gap, we present the first real-world BEV 3D detection benchmark that integrates both pinhole and fisheye cameras, constructed by converting KITTI-360 data into the nuScenes format. We systematically evaluate strategies including image rectification, MEI camera model–guided distortion-aware view transformation, and polar coordinate representations. Our experiments demonstrate that projection-free architectures significantly outperform conventional view transformation approaches under fisheye distortion, highlighting their superior robustness and offering practical design guidelines for cost-effective, robust 3D perception systems in autonomous driving.

BEV object detectionfisheye camerasmulti-view perception

This work addresses the limitations of existing bird’s-eye-view (BEV) perception methods, which suffer from constrained geometric accuracy and semantic consistency due to the absence of explicit 3D geometric modeling. To overcome this, the paper introduces 3D Gaussian Splatting reconstruction into the BEV perception framework for the first time, leveraging multi-view images to explicitly construct a high-fidelity 3D scene representation and generate geometrically aligned BEV features. By effectively integrating explicit 3D geometry with semantic information, the proposed approach significantly enhances model interpretability and perception performance. State-of-the-art results on the nuScenes and Argoverse benchmarks demonstrate the efficacy and potential of explicit 3D reconstruction in advancing BEV-based perception systems.

3D geometric understandingautonomous drivingBEV perception

This work addresses the significant performance degradation in cross-dataset radar-camera bird’s-eye-view (BEV) perception caused by discrepancies in scenes, sensor configurations, and environmental conditions. Existing approaches fail to effectively model the feature variations originating from the source domain. To tackle this, the paper introduces frequency-domain scene variation modeling into multimodal 3D perception for the first time. By synthesizing diverse source-domain views through frequency-domain transformations, it analyzes their impact on BEV fusion features and proposes a spatial regularization mechanism integrated within the fusion process. This approach enhances model generalization without requiring any target-domain data. Consistent performance gains are demonstrated across multiple BEV backbones on cross-dataset 3D detection tasks between View-of-Delft and TJ4DRadSet, with notable improvements persisting even when limited target-domain data is available.

3D object detectionBEV perceptioncross-dataset generalization

Hot Scholars

XC

Xieyuanli Chen

Associate Professor, NUDT, China
RoboticsSLAMLocalizationLiDAR Perception
KJ

Kun Jiang

Tsinghua University
autonomous driving
GS

Ganesh Sistu

Principal Artificial Intelligence Architect, Valeo Ireland
Autonomous DrivingMachine LearningComputer VisionDeep Learning
SX

Shaoyuan Xie

University of California, Irvine
AI SecurityMachine LearningAutonomous Driving