🤖 AI Summary
Existing global building footprint raster products lack independent, comprehensive, and equitable accuracy assessments. This study presents the first systematic evaluation of four major products—GHSL, TEMPO, GBA, and Overture—using the newly developed, human-annotated global benchmark dataset ORBITaL-Net under a unified framework across multiple spatial resolutions. Accuracy is analyzed stratified by geographic region, population density, and national income level, employing standard remote sensing metrics and multidimensional statistical methods. Results reveal that while GBA or TEMPO generally perform best overall, all products exhibit substantially degraded accuracy in Africa, Asia, and high-density urban areas. This work establishes a fine-grained, reproducible evaluation framework to inform the selection and improvement of global building mapping products.
📝 Abstract
Geo-spatial rasters of building footprint area are useful for a variety of tasks, such as monitoring urbanization, improving energy efficiency, and tracking greenhouse gas emissions. There are now multiple global building raster datasets, however there lacks an independent, comprehensive, and fair assessment of their accuracy. In this work, we evaluate the accuracy of four major global building products: Global Human Settlement Layer (GHSL), Microsoft's TEMPO (TEMPO), The Global Building Atlas (GBA), and Overture. As ground truth for assessing their accuracy, we use ORBITaL-Net, a globally diverse dataset of manually labeled building footprints. To ensure fairness, we evaluate products on grids of multiple spatial resolutions, and several conventional performance metrics. Our results indicate that either GBA or TEMPO generally achieves the highest overall accuracy, depending upon the particular evaluation criteria. We also stratify the accuracy of each product by several factors: geographic location, population density, and income groups. The results reveal that product accuracy can sometimes vary significantly with respect to these factors. Notably, all products are significantly less accurate in Africa and Asia. Most products also suffer significant accuracy reduction in high-density urban areas.