🤖 AI Summary
This study addresses the noise and void artifacts prevalent in stereo photogrammetric digital surface models (DSMs) by proposing a multimodal elevation inpainting method based on an improved Stable Diffusion 3 architecture. The framework conditions generative modeling on both photogrammetric DSMs and satellite imagery, incorporating a pruned text stream and chunk-wise normalization strategy to facilitate stable transfer from the natural image domain to the elevation domain. Validated against LiDAR ground truth, the proposed approach substantially enhances DSM accuracy, reducing the root mean square error (RMSE) from 6.00 to 3.45 meters in dense urban environments and from 4.16 to 2.77 meters in cross-city generalization tests. These results demonstrate that the method offers an effective solution for acquiring high-precision elevation data at low cost.
📝 Abstract
Large-scale Digital Surface Models (DSMs) can be produced cost-effectively from satellite images via stereo-photogrammetry. However, the resulting 3D maps are often contaminated by noise, outliers, and voids. On the other hand, aerial LiDAR provides high-accuracy elevation measurements at a substantially higher cost. In this work, we study diffusion models conditioned both on photogrammetric DSMs and Pléiades imagery to refine vertically co-registered DSMs. We introduce a modified Stable Diffusion 3 architecture with a pruned text stream and a patch-wise normalization strategy, enabling stable training on LiDAR data and transfer from natural images to elevation maps. Experiments in French cities demonstrate that multimodal conditioning improves elevation accuracy, reducing Dense Urban RMSE from 6.00 to 3.45 m in the in-context cities and from 4.16 to 2.77 m in the held-out city of Bordeaux.