π€ AI Summary
Existing methods for 3D building reconstruction from single faΓ§ade images suffer from oversimplified geometry, homogeneous appearance, and severe structural ambiguity. To address these challenges, this paper proposes a map-contour-driven, three-stage high-fidelity 3D building generation framework: (1) height map estimation from vectorized map contours, (2) geometric reconstruction, and (3) appearance stylization. Our method innovatively integrates a customized ControlNet with Text2Mesh to enable joint controllable generation of geometry and texture. By employing multi-stage generative modeling, it explicitly decouples structural and appearance representation. Extensive evaluation on real-world urban planning maps and geospatial data demonstrates that our approach produces geometrically accurate, detail-rich 3D building models. It significantly reduces manual modeling effort and enhances efficiency for city-scale 3D reconstruction.
π Abstract
We introduce GeoTexBuild, a modular generative framework for creating 3D building models from map footprints. The proposed framework employs a three-stage process comprising height map generation, geometry reconstruction, and appearance stylization, culminating in building models with intricate geometry and appearance attributes. By integrating customized ControlNet and Text2Mesh models, we explore effective methods for controlling both geometric and visual attributes during the generation process. By this, we eliminate the problem of structural variations behind a single facade photo of the existing 3D generation techniques. Experimental results at each stage validate the capability of GeoTexBuild to generate detailed and accurate building models from footprints derived from site planning or map designs. Our framework significantly reduces manual labor in modeling buildings and can offer inspiration for designers.