🤖 AI Summary
To address instance incompleteness and geometric shape mismatch in vectorized high-definition map construction from surround-view imagery—caused by frontal-view feature loss—this paper proposes a dual-path feature enhancement framework. First, we design a novel dual-enhancement module integrating explicit fusion and implicit modulation to enable efficient multi-view feature collaboration. Second, we introduce frontal-view keypoint supervision to guide BEV feature learning with geometric structural priors. Third, we develop an end-to-end BEV vectorization decoder. Evaluated on nuScenes and OpenLane benchmarks, our method achieves state-of-the-art performance, significantly improving instance completeness and geometric accuracy for lane markings, curbs, and other road elements—yielding +3.2% in F-score and +2.8% in AP₅₀.
📝 Abstract
Constructing vectorized high-definition maps from surround-view cameras has garnered significant attention in recent years. However, the commonly employed multi-stage sequential workflow in prevailing approaches often leads to the loss of early-stage information, particularly in perspective-view features. Usually, such loss is observed as an instance missing or shape mismatching in the final birds-eye-view predictions. To address this concern, we propose a novel approach, namely extbf{HybriMap}, which effectively exploits clues from hybrid features to ensure the delivery of valuable information. Specifically, we design the Dual Enhancement Module, to enable both explicit integration and implicit modification under the guidance of hybrid features. Additionally, the perspective keypoints are utilized as supervision, further directing the feature enhancement process. Extensive experiments conducted on existing benchmarks have demonstrated the state-of-the-art performance of our proposed approach.