Score
Design, build, and train perception models and processing pipelines that convert multi-view images, point clouds, or other sensor inputs into bird’s‑eye‑view (BEV) representations and maps, including object and semantic layouts in that top‑down coordinate frame. Analyze and evaluate BEV representations for spatial alignment, fidelity, temporal consistency, and sensor‑fusion strategies used to produce them.
BEV perception faces reliability bottlenecks in safety-critical scenarios—such as occlusion, adverse weather, and dynamic traffic—hindering its deployment in autonomous driving. Method: This paper presents the first systematic survey of BEV perception evolution from a functional safety perspective, categorizing advancements into three phases: unimodal, multimodal onboard, and multi-agent collaborative perception. It proposes a unified taxonomy for open-world challenges, identifying core issues including sensor degradation, unknown-class recognition, and low-latency coordination. The survey integrates multi-sensor fusion, open-set recognition, label-free learning, degradation-resilient modeling, and vehicle–road–cloud low-latency communication, while reviewing mainstream architectures and benchmark datasets. Contribution/Results: It clarifies critical technical bottlenecks and provides theoretical foundations and practical pathways for advancing BEV perception toward end-to-end autonomous driving, embodied intelligence, and large-model-augmented systems.
This work addresses the limitations of existing bird’s-eye-view (BEV) perception methods, which suffer from constrained geometric accuracy and semantic consistency due to the absence of explicit 3D geometric modeling. To overcome this, the paper introduces 3D Gaussian Splatting reconstruction into the BEV perception framework for the first time, leveraging multi-view images to explicitly construct a high-fidelity 3D scene representation and generate geometrically aligned BEV features. By effectively integrating explicit 3D geometry with semantic information, the proposed approach significantly enhances model interpretability and perception performance. State-of-the-art results on the nuScenes and Argoverse benchmarks demonstrate the efficacy and potential of explicit 3D reconstruction in advancing BEV-based perception systems.
In multi-view pure-vision 3D perception, dynamic objects impede temporal alignment of bird’s-eye-view (BEV) features across frames. To address this, we propose the Motion-guided BEV Fusion Network (MBFNet), which introduces a recurrent online temporal fusion architecture that explicitly models object motion and achieves spatiotemporal alignment of BEV features across frames. MBFNet incorporates a motion-guided dynamic alignment mechanism and an end-to-end differentiable temporal consistency loss to enhance fusion quality between historical and current BEV features. Crucially, the method requires no auxiliary motion annotations and is trained end-to-end solely from image inputs. Evaluated on the nuScenes benchmark, MBFNet achieves 63.9% NDS—the highest performance among pure-vision 3D detectors at the time—demonstrating for the first time the critical role of efficient online temporal fusion in learning robust dynamic BEV representations.
To address misalignment of BEV features and depth estimation errors caused by LiDAR-camera calibration inaccuracies in multimodal 3D object detection, this paper proposes a dual-module correction framework: Local Align and Global Align. Local Align performs neighborhood-aware, graph-matching-based self-correction of depth estimates, explicitly modeling local geometric mismatches. Global Align mitigates global projection distortions by optimizing cross-modal alignment in the BEV feature space. This work is the first to explicitly model and compensate for geometric mismatches induced by calibration noise within the BEV fusion paradigm. On the nuScenes validation set, our method achieves a 70.1% mAP—outperforming BEV Fusion by 1.6%. Under injected calibration noise, the performance gain widens to 8.3%, demonstrating significantly enhanced robustness. The framework thus bridges a critical gap between theoretical calibration assumptions and practical sensor deployment, improving both accuracy and reliability in real-world multimodal 3D perception.
To address the insufficient robustness of bird’s-eye-view (BEV) perception in autonomous driving under atypical and challenging conditions—such as camera degradation and multi-sensor failures—this paper introduces RoboBEV, the first out-of-distribution (OOD) evaluation benchmark for BEV perception covering four tasks: 3D detection, semantic segmentation, depth estimation, and occupancy prediction, with systematic evaluation of 33 models. We propose a multi-granularity BEV robustness evaluation framework, revealing for the first time the weak correlation between standard performance and robustness. Additionally, we design a CLIP-based temporal enhancement strategy and a depth-agnostic BEV transformation method. Experiments demonstrate that pretraining and depth-agnostic BEV representations significantly improve robustness; temporal fusion boosts average robustness by 12.6%; and our CLIP-based strategy reduces mAP degradation by 37% under severe sensor degradation.
To address feature distortion and boundary ambiguity in multi-camera BEV 3D detection—stemming from the ill-posed image-to-3D mapping—this paper proposes an object-aware pseudo-3D–depth joint modeling framework. Our method introduces two key innovations: (1) a novel object-level depth supervision scheme coupled with pseudo-voxel encoding, enhancing structural consistency in depth estimation; and (2) a 2D detection-guided foreground pixel projection mechanism integrated with deformable attention fusion, enabling precise spatial alignment and contextual enhancement. These components collectively improve object structural representation and spatial localization accuracy in the BEV feature space. Evaluated on nuScenes, our approach achieves 72.5% mAP and 78.3% NDS, outperforming state-of-the-art baselines including BEVDet and BEVFormer.
This work addresses the performance degradation of existing BEV-based 3D object detection methods—primarily designed for pinhole cameras—when applied to fisheye cameras due to severe radial distortion, and the absence of benchmarks or effective solutions for hybrid camera setups. To bridge this gap, we present the first real-world BEV 3D detection benchmark that integrates both pinhole and fisheye cameras, constructed by converting KITTI-360 data into the nuScenes format. We systematically evaluate strategies including image rectification, MEI camera model–guided distortion-aware view transformation, and polar coordinate representations. Our experiments demonstrate that projection-free architectures significantly outperform conventional view transformation approaches under fisheye distortion, highlighting their superior robustness and offering practical design guidelines for cost-effective, robust 3D perception systems in autonomous driving.
To address challenges in infrastructure-side multi-camera 3D object detection—including multi-view geometric heterogeneity, diverse camera configurations, degraded visual quality, and complex road layouts—this paper proposes a Transformer-based bird’s-eye view (BEV) perception framework. The framework supports flexible integration of heterogeneous cameras and introduces a graph-enhanced fusion module that explicitly models camera-to-BEV-grid geometric relationships while jointly aggregating implicit visual features for relation-aware multi-view feature fusion. It further incorporates deformable attention, graph neural networks, and multimodal fusion. Extensive experiments on the synthetic dataset M2I and the real-world dataset RoScenes demonstrate superior performance. Notably, the method maintains high accuracy under adverse conditions such as extreme weather and sensor degradation, achieving state-of-the-art results on both benchmarks and exhibiting strong potential for deployment in practical intelligent transportation systems.
This work addresses the challenges of multi-camera data fusion and insufficient bird’s-eye-view (BEV) segmentation accuracy in complex driving scenarios by proposing a Transformer-based Variational Bifurcated network (TVB). TVB is the first approach to integrate variational inference with normalizing flows for BEV segmentation. It implicitly learns the mapping from multi-view images to a unified BEV representation through posterior BEV supervision, generating multiple candidate maps. A novel BEV-attention fusion module adaptively aggregates these candidates to enhance the realism and expressiveness of the resulting map. Evaluated on the nuScenes and OPV2V datasets, the proposed method significantly outperforms existing approaches in both multi-camera BEV segmentation and lane perception tasks, achieving state-of-the-art performance.
This work addresses the challenging problem of cross-modal calibration between LiDAR and cameras by proposing the first bird’s-eye-view (BEV)-based alignment framework. The method unifies multimodal data into a shared BEV space and employs a two-stage optimization strategy: it first implicitly regresses coarse calibration parameters and then explicitly aligns cross-modal features, enhanced by a CLIP-inspired contrastive loss to enforce semantic consistency. By integrating domain-specific BEV feature extraction with contrastive learning constraints, the approach significantly outperforms existing methods, achieving state-of-the-art calibration accuracy. On the KITTI and nuScenes benchmarks, it reduces relative rotation error by 51% and 68%, and translation error by 80% and 91%, respectively.