Score
Designs and implements pipelines, algorithms, and tools to ingest sensor or survey data and produce, align, compress, update, and validate high-definition (HD) maps — detailed geometric and semantic representations of environments. Evaluates and optimizes map accuracy, consistency, storage/streaming formats, and interfaces for downstream functions such as precise localization, perception fusion, and motion planning.
Autonomous driving maps face concurrent challenges in achieving high precision, lightweight representation, and end-to-end integration. To address these, this paper proposes a three-stage evolutionary paradigm: High-Definition (HD) maps, Lite maps, and Implicit maps—systematically analyzing their representational forms, production pipelines, and fundamental bottlenecks. We introduce the first unified taxonomy integrating semantic compression, neural radiance fields (NeRF), differentiable rendering, and end-to-end learning to enable a paradigm shift from explicit geometric storage to implicit neural representation. Our contributions include a full-lifecycle technical roadmap, a multi-stage collaborative mapping framework, and an automated production pipeline. These advances significantly improve map generalizability and deployment efficiency. The work establishes both theoretical foundations and practical guidelines for next-generation autonomous driving map research and development.
This work addresses the challenge of high-definition (HD) map generation in regions lacking professional surveying infrastructure, where conventional approaches rely on costly dense sensor suites and high-precision reference data. The authors propose a lane-level HD map construction pipeline that operates solely on publicly available geospatial engineering data and adopts a lanelet-based representation. Notably, they introduce a constraint-driven validation mechanism that requires no external reference, enabling self-consistency checks through geometric, topological, and elevation-based regulatory constraints. This approach substantially enhances the modularity and auditability of the mapping workflow. Evaluated on real-world road networks across four cities in Lower Saxony, Germany, the method demonstrates robust performance, achieving a 100% defect detection rate with zero false positives in controlled defect-injection experiments.
Autonomous vehicles face dual challenges in online high-definition (HD) map construction: limited onboard sensor perception range and high HD map maintenance costs. This paper proposes, for the first time, leveraging open-text-annotated low-definition (LD) maps—such as OpenStreetMap—as prior knowledge to enable real-time, long-range, high-accuracy local HD map generation. Methodologically, we introduce a text-semantics-augmented SD (semantic-detail) map representation, a point-level SD map encoder, and an orthogonal element identification mechanism to jointly model heterogeneous map elements without relying on predefined categories. Furthermore, we design a multimodal alignment and fusion neural network integrating NLP feature embeddings, vector-map token enhancement, and point-level graph encoding. Evaluated on Argoverse 2 and nuScenes, our approach achieves a +5.9 mAP improvement (+45%) over prior-free methods and a +3.2 mAP gain (+20%) over existing SD-prior-based approaches.
To address the high construction cost, infrequent updates, and poor sensor generalization of conventional high-definition (HD) maps in dynamic campus environments, this paper proposes a real-time online mapping method leveraging tightly coupled stereo camera and LiDAR fusion. The core contribution is a lightweight semantic vector map (SemVecMap) framework, integrating multi-sensor spatiotemporal calibration, 3D semantic segmentation, and incremental map updating—fine-tuned on campus-specific data to enhance generalization. Evaluated on a real-world golf-cart autonomous platform, the system achieves sub-meter-accurate, real-time 3D HD map generation with adaptive updates under dynamic conditions. It significantly reduces manual intervention, improves map freshness and scene adaptability, and delivers a practical, lightweight mapping solution tailored for autonomous driving in confined, semi-structured environments such as university campuses.
Existing high-definition (HD) map datasets are limited in scale, semantic richness, and multimodal support, hindering long-horizon autonomous driving map construction. This work proposes HRDX, a large-scale vectorized HD map dataset spanning 1,400 kilometers of roadways, integrating six-camera imagery, 128-beam LiDAR, RTK GNSS/IMU, and precisely aligned aerial orthophotos. It annotates ten map element classes and over twenty semantic topological attributes. Aerial imagery is innovatively leveraged as a structural prior, and a composite scoring metric is introduced to jointly evaluate geometric and attribute accuracy. The dataset supports multimodal bird’s-eye-view (BEV) fusion and learning with privileged information. Experiments demonstrate that HRDX substantially improves online vector map construction performance, with aerial imagery enhancing geometric fidelity and enabling knowledge distillation to transfer these gains to purely vision-based models.
Online high-definition (HD) maps for autonomous driving suffer from temporal instability caused by sensor pose drift, yet existing work focuses solely on per-frame accuracy while neglecting cross-frame stability. To address this gap, we introduce the first comprehensive benchmark dedicated to evaluating temporal stability of online HD maps. We propose a multi-dimensional stability assessment framework, defining three novel metrics—existence stability, localization stability, and shape stability—and a unified mean Average Stability (mAS) score. Extensive experiments across 42 models and variants demonstrate that accuracy (measured by mAP) and stability (mAS) constitute largely orthogonal performance dimensions. Our analysis identifies key architectural and training design factors—such as temporal consistency regularization, pose-aware feature alignment, and stability-augmented loss functions—that jointly enhance both mAP and mAS. These findings provide both theoretical insights and practical guidelines for developing more robust online HD mapping systems.
This study addresses the lack of systematic best practices in large-scale Earth observation (EO) mapping, which often introduces errors during data preprocessing, model training, inference deployment, and validation, thereby compromising the reliability and scientific credibility of map products. To remedy this, we propose the first end-to-end best practice framework for EO mapping, encompassing the entire workflow from satellite data acquisition to operational map delivery. The framework integrates six core components: EO data infrastructure, preprocessing, machine learning dataset construction, uncertainty quantification, map production and dissemination, and independent validation. Emphasizing the interdependence of these stages, it embeds uncertainty quantification and independent validation as integral elements. By synergizing machine learning, distributed computing, and geospatial validation techniques, the framework establishes a reproducible and scalable mapping pipeline that substantially enhances the quality, consistency, and scientific rigor of EO-derived maps, supported by open-source resources to foster community adoption.
Existing approaches struggle to effectively align and fuse multimodal data—such as multi-view images, semantic maps, and satellite imagery—limiting the accuracy and robustness of online high-definition (HD) map construction. Inspired by human driving cognition, this work proposes Driver2Map, the first model to enable collaborative mapping across these three modalities. It employs a two-stage alignment strategy to mitigate cross-modal spatial misalignment, introduces a pose-guided bird’s-eye-view (BEV) fusion module to handle dynamic occlusions, and incorporates a map refinement mechanism leveraging pre-trained priors to enhance structural and textural consistency. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art techniques in terms of Intersection over Union (IoU) and Average Precision (AP), achieving more accurate and robust online HD map generation.
Existing high-definition map construction methods struggle to simultaneously achieve geometric accuracy and topological correctness: vectorization-based approaches preserve structural integrity but suffer from geometric distortions, whereas rasterization-based methods offer precise geometry yet lack explicit structural representation. To address this limitation, this work proposes GSMap, a novel framework that introduces learnable 2D Gaussian sequences to represent map elements, modeling vector vertices as Gaussian centers. By integrating differentiable rasterization for pixel-level geometric constraints and topology-aware vectorization to enforce structural regularity, GSMap enables end-to-end joint optimization of geometry and topology. The method significantly outperforms existing approaches on both nuScenes and Argoverse2 benchmarks while remaining compatible with mainstream HD map architectures.