Score
Designs and implements methods and pipelines to detect, segment, and vectorize building outlines (footprints) from spatial raster data such as satellite or aerial imagery and derived maps, producing georeferenced polygon outputs. Builds tools to align and compare those footprints across time or datasets to support tasks like bi‑temporal registration and structural change detection.
This study addresses the prevalence of topological errors in building footprints generated by deep learning models, which hinder their direct integration into GIS databases. To tackle this issue, the authors propose a multi-domain GeoAI quality control framework that fuses 24-dimensional features encompassing geometric, spatial contextual, and spectral-textural attributes. By integrating geometric regularization and spatial mutual exclusion constraints, the framework enables object-level automated quality inspection and purification. Initial masks are produced using U-Net (ResNet-34) and SAM-LoRA (ViT-B), followed by boundary deformation and duplicate object detection via decision tree classifiers. Experimental results demonstrate that the framework achieves 95.31% accuracy, 91.06% F1-score, and 0.880 Matthews correlation coefficient on an independent test area, with an 87.34% error footprint detection rate. This approach significantly enhances building database purity to 95.38% and reduces relative error by 83.09%, offering a robust and transferable quality assurance mechanism for automated GIS production.
Optical imagery-based building footprint extraction is often compromised by occlusion, perspective distortion, and the absence of elevation data, leading to incomplete or misaligned results. This work proposes the first large-scale vectorized building footprint dataset and benchmark specifically designed for airborne LiDAR point clouds, encompassing 33,000 urban and rural tiles of size 128×128 meters, along with 3,000 cross-domain test tiles to evaluate geographic generalization. The dataset provides precisely aligned vector footprints coupled with elevation information, enabling fine-grained modeling and cross-regional studies. Through comprehensive baseline experiments, this study highlights key challenges including high intra-class variability, data imbalance, and noise, thereby advancing research in urban perception and building modeling.
Existing methods for 3D building reconstruction from single façade images suffer from oversimplified geometry, homogeneous appearance, and severe structural ambiguity. To address these challenges, this paper proposes a map-contour-driven, three-stage high-fidelity 3D building generation framework: (1) height map estimation from vectorized map contours, (2) geometric reconstruction, and (3) appearance stylization. Our method innovatively integrates a customized ControlNet with Text2Mesh to enable joint controllable generation of geometry and texture. By employing multi-stage generative modeling, it explicitly decouples structural and appearance representation. Extensive evaluation on real-world urban planning maps and geospatial data demonstrates that our approach produces geometrically accurate, detail-rich 3D building models. It significantly reduces manual modeling effort and enhances efficiency for city-scale 3D reconstruction.
This paper addresses the task of converting building instance masks in remote sensing imagery into vectorized polygon representations. We propose the Global Collinearity-aware Polygonization (GCP) framework, which takes binary instance segmentation masks as input and employs a Transformer-based regression model to predict polygonal contour polylines. A differentiable collinearity-aware objective function is designed and integrated with dynamic programming to enable end-to-end training for globally optimal polygon simplification. Key contributions include: (i) the first explicit incorporation of collinearity modeling into the simplification objective with joint optimization; (ii) elimination of local heuristic strategies, thereby guaranteeing theoretically sound global optimality; and (iii) no reliance on prior geometric assumptions, ensuring strong generalizability. GCP achieves state-of-the-art performance on two major public benchmarks, significantly outperforming classical methods such as Douglas–Peucker. The source code is publicly available.
To address the weak generalization and high human-in-the-loop cost in extracting unseen building footprints from ultra-high-resolution aerial imagery, this paper proposes a promptable end-to-end framework that jointly predicts roof segmentation and instance-level roof-to-footprint offset vectors. Key contributions include: (i) the first Offset-Building Model (OBM), explicitly modeling geometric displacement between roofs and footprints; (ii) distance-aware non-maximum suppression (DNMS) to enhance detection robustness under scale and density variations; and (iii) a prompt-based evaluation paradigm enabling efficient human–AI collaboration. Evaluated on the Huizhou test set (7,000+ samples), our method significantly outperforms baselines: offset error decreases by 16.6%, roof IoU improves by 10.8%, and offset vector loss drops by 6.5%. These results demonstrate superior generalization capability and practical engineering applicability for large-scale footprint extraction.
Existing spatial pattern matching methods are largely confined to two-dimensional space and struggle to handle three-dimensional entity matching involving elevation or height information in real-world scenarios. This work extends spatial pattern matching to 3D environments for the first time, introducing a general problem formulation and proposing a subgraph-matching-based algorithm that explicitly models distance relationships in three-dimensional space. To support empirical evaluation, the authors construct the first 3D spatial pattern matching dataset, integrating both synthetic data and real-world building structures from the city of Hamburg. Experimental results on this benchmark demonstrate the effectiveness of the proposed approach, establishing a foundational algorithmic framework and experimental platform for future research in 3D spatial pattern analysis.
This work addresses the challenge of building footprint extraction from high-resolution remote sensing imagery, where complex structures and varying imaging conditions often degrade performance. Conventional approaches typically rely on multi-stage post-processing pipelines, leading to low efficiency and error propagation. To overcome these limitations, the authors propose PolyBuild, an end-to-end model that directly generates vectorized building polygons without any post-processing—a first in the field. PolyBuild integrates a CNN-Transformer hybrid architecture: an initial contour generation module jointly performs detection and coarse outline extraction, while a Transformer-based decoder refines the contour by effectively fusing local details with global contextual information. Extensive experiments on three benchmark building datasets demonstrate that PolyBuild significantly outperforms state-of-the-art mask-based and contour-based methods, confirming its superior accuracy and robustness.
This work addresses the limitation of existing remote sensing change detection methods, which typically assume a single change event per location and thus fail to capture the complex, recurrent dynamics of urban building changes. To this end, we introduce the Urban Building Dynamics Detection task, which models building changes through dynamic footprints—temporal intervals representing when changes occur—and provides a unified representation for both single and multiple change events. We propose FootprintNet, a novel architecture incorporating a state-action interaction mechanism and state transition constraints, leveraging temporal boundary cues to enable pixel-level dynamic footprint classification. Furthermore, we present a new evaluation metric, the Building Change Dynamics Score (BCDS), which jointly assesses semantic correctness and temporal alignment. Experiments on the TSCD, MUDS, and WUSU datasets demonstrate that our method significantly outperforms existing approaches in multi-temporal building change detection.
Existing global building footprint raster products lack independent, comprehensive, and equitable accuracy assessments. This study presents the first systematic evaluation of four major products—GHSL, TEMPO, GBA, and Overture—using the newly developed, human-annotated global benchmark dataset ORBITaL-Net under a unified framework across multiple spatial resolutions. Accuracy is analyzed stratified by geographic region, population density, and national income level, employing standard remote sensing metrics and multidimensional statistical methods. Results reveal that while GBA or TEMPO generally perform best overall, all products exhibit substantially degraded accuracy in Africa, Asia, and high-density urban areas. This work establishes a fine-grained, reproducible evaluation framework to inform the selection and improvement of global building mapping products.
This study addresses the spatial distortion in oblique-view remote sensing imagery caused by geometric projection displacement between building rooftops and their ground footprints. To this end, it formulates roof-to-footprint offset vector (RFOV) extraction as an independent learning task, decoupling geometric correction from semantic segmentation. The authors introduce ObliCity, the first large-scale oblique urban dataset, which integrates high-resolution drone and satellite imagery, and propose DragRoof—a framework based on ordinary differential equations (ODEs) that simulates a continuous, annotation-inspired dragging process to adaptively learn geometrically consistent offset fields. Experiments demonstrate that the proposed method achieves state-of-the-art RFOV extraction accuracy on ObliCity with fewer inference steps, significantly outperforming existing approaches in both direction and magnitude estimation.