🤖 AI Summary
This study addresses the challenge of generating B-reps from single images while simultaneously ensuring geometric accuracy, topological validity, and shape complexity. We propose a geometry-first framework that leverages pretrained models to generate feature grids as a unified intermediate representation. A dual-decoder architecture subsequently extracts geometric and topological cues in a decoupled manner. Topology is then recovered via grid regions, overcoming predefined face-count limitations and enabling complexity-adaptive scaling, before explicit B-reps are assembled using a CAD kernel. Experiments on the DeepCAD benchmark demonstrate that our method achieves an 80.49% validity rate and significantly reduces face Chamfer distance, outperforming CADDreamer and HoLa while exhibiting superior generalization to high-complexity shapes.
📝 Abstract
Generating a boundary representation (B-rep) conditioned on a single image requires faithful reconstruction of geometry, valid topology, and support for complex shapes. We present UniBRep, a geometry-first framework that adapts a pretrained image-to-3D model to generate a feature mesh as a unified intermediate representation. Its surface provides a geometric scaffold, while spatially aligned learned features encode face-separation cues for topology recovery. Dual decoder branches generate the geometry and face-separation features; a geometry- and feature-guided construction pipeline then fits parametric surfaces, recovers boundary curves and connectivity, and assembles an explicit B-rep using a CAD kernel. Recovering topology from mesh regions avoids predefined architectural face-count limits, allowing face count to scale with shape complexity. On the standard DeepCAD benchmark, UniBRep produces valid B-reps for 80.49\% of inputs and reduces face Chamfer distance from 0.1096 to 0.0345 relative to CADDreamer. In a matched comparison, UniBRep also outperforms the HoLa public demo across all reported metrics. Further evaluations demonstrate scalability to high-complexity shapes beyond the standard 30-face range, generalization to objects outside the CAD training distribution, and qualitative transfer to real photographs.