🤖 AI Summary
Existing HD map construction methods suffer from geometric misalignment and poor generalization in image-to-bird’s-eye-view (BEV) projection, often generating spurious road elements that degrade vectorization accuracy. To address this, we propose a BEV mapping framework integrating camera-geometric constraints with scene-adaptive fusion. First, we design a probabilistic projection module grounded in camera parameters and confidence scoring to suppress hallucination. Second, we introduce geometric-prior-guided mapping optimization, coupled with a confidence-weighted temporal attention mechanism for robust feature accumulation across frames. Evaluated on newly partitioned nuScenes and Argoverse2 benchmarks, our method significantly improves road structure reconstruction accuracy—especially at long ranges and in complex scenes—achieving a 5.2% mAP gain over prior approaches. This demonstrates superior generalization capability and practical effectiveness for real-world HD mapping.
📝 Abstract
Constructing high-definition (HD) maps from sensory input requires accurately mapping the road elements in image space to the Bird's Eye View (BEV) space. The precision of this mapping directly impacts the quality of the final vectorized HD map. Existing HD mapping approaches outsource the projection to standard mapping techniques, such as attention-based ones. However, these methods struggle with accuracy due to generalization problems, often hallucinating non-existent road elements. Our key idea is to start with a geometric mapping based on camera parameters and adapt it to the scene to extract relevant map information from camera images. To implement this, we propose a novel probabilistic projection mechanism with confidence scores to (i) refine the mapping to better align with the scene and (ii) filter out irrelevant elements that should not influence HD map generation. In addition, we improve temporal processing by using confidence scores to selectively accumulate reliable information over time. Experiments on new splits of the nuScenes and Argoverse2 datasets demonstrate improved performance over state-of-the-art approaches, indicating better generalization. The improvements are particularly pronounced on nuScenes and in the challenging long perception range. Our code and model checkpoints are available at https://github.com/Fatih-Erdogan/mapping-like-skeptic .