Look Where You Can: Active View Selection for CAD Reconstruction under Occlusion

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the accuracy degradation in CAD reconstruction under occluded scenarios caused by limited viewpoints by proposing the SightCAD framework. This work pioneers a joint training paradigm for viewpoint selection and CAD generation, departing from the conventional full-viewpoint coverage assumption by treating viewpoint feasibility as a constraint. By leveraging reinforcement learning to co-optimize a selector and a vision-language model (VLM), the framework achieves a closed loop between parametric CAD code generation and geometric evaluation. Experimental results demonstrate that the proposed method improves the effective mIoU by 21 percentage points over baselines while achieving state-of-the-art code executability, substantially enhancing 3D reconstruction quality in occluded environments.
📝 Abstract
CAD reconstruction methods assume a luxury reality rarely grants: unrestricted visual access to the object, photographed from any desired angle. Real objects, however, are scene-embedded, bolted against walls, wedged into corners, resting on floors, where the scene renders much of the view sphere unreachable and the remaining views unequally informative. We introduce \textbf{SightCAD}, a framework for parametric CAD reconstruction that treats view feasibility as a first-class constraint. In this work we consider objects from standard CAD benchmarks embedded in realistic indoor scenes with physically derived visibility constraints over a discrete view sphere. A learned view selector must choose $K$ feasible views for a vision--language model (VLM) that generates executable CadQuery code, scored by geometric fidelity of the executed solid. Because reward arrives only after discrete view selection, autoregressive generation, and CAD-kernel execution, we propose a joint training paradigm in which the view selector and the CAD-generation VLM are trained together against this reward. The learned selection policy departs sharply from random, uniform, and coverage-greedy alternatives, outperforming surface-area maximization (SA-max) by up to $6.4$ Intersection-over-Union (IoU) points across budgets $K\in\{1,\dots,5\}$. The full system surpasses strong external baselines on scene-embedded, occluded multi-view renders of DeepCAD and Fusion360 objects ($+21$ and $+17$ effective-mIoU points over the best baseline, respectively), as well as on test-time domain-canonicalized real images from the industrial T-LESS benchmark and on both synthetic and real images from the MP6D industrial metal-parts benchmark, while producing the highest rate of executable programs of any method compared (invalid-code rate ${\leq}1.5\%$).
Problem

Research questions and friction points this paper is trying to address.

CAD reconstruction
occlusion
active view selection
scene-embedded objects
visibility constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

Active View Selection
Parametric CAD Reconstruction
Joint Training Paradigm
Vision-Language Model
Occlusion Handling
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kartik Bali
Helmholtz-Zentrum Hereon
M
Mahish Guru
Leuphana Universität Lüneburg
Y
Yiderigun Borjigin
Universität des Saarlandes
A
Alexandra Starostina
Helmholtz-Zentrum Hereon
Christian J. Cyron
Christian J. Cyron
Professor at Hamburg University of Technology, Germany
solid mechanics - computational mechanics - materials modeling - micromechanics
Roland Aydin
Roland Aydin
Professor at Hamburg University of Technology, Germany
Large Language ModelsMachine LearningMaterials Science