VGGT-CAD: Reconstructing Parametric CAD 3D Model with Geometric Grounding

๐Ÿ“… 2026-09-17
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๆๅ‡บVGGT-CAD๏ผŒ้€š่ฟ‡ๅ‡ ไฝ•ๅ…ˆ้ชŒๅ’Œๅคš่ง†ๅ›พ็‰นๅพ่žๅˆ่งฃๅ†ณๅ‚ๆ•ฐๅŒ–CADๆจกๅž‹ไปŽๆœ‰้™่ง†่ง’้‡ๅปบ็š„้—ฎ้ข˜ใ€‚
๐Ÿ“ Abstract
Parametric CAD reconstruction requires recovering both precise geometry and editable modeling operations from visual observations, making it challenging under limited and ambiguous views. Existing methods mainly rely on 2D appearance cues and lack strong multi-view geometric priors. In this work, we present VGGT-CAD, a geometry-aware framework for parametric CAD reconstruction from single- and multi-view observations. We transfer pretrained 3D geometric priors into CAD reconstruction by encoding camera parameters as condition tokens and jointly modeling them with image tokens. To handle varying numbers of viewpoints, we introduce a variable-view cross-view context aggregation module that adaptively fuses multi-view features. We further develop a training-free geometry-aware view selection strategy to select complementary and reliable frames during inference. The resulting representation is decoded into CAD command sequences using a non-autoregressive decoder. We also develop VideoCAD, a large-scale multi-view video benchmark derived from existing CAD data through multi-view re-rendering. Extensive experiments demonstrate the effectiveness of VGGT-CAD for visual CAD reconstruction under different observation configurations.
Problem

Research questions and friction points this paper is trying to address.

parametric CAD
reconstruction
visual observations
geometric priors
multi-view
Innovation

Methods, ideas, or system contributions that make the work stand out.

geometry-aware framework
variable-view cross-view context aggregation
non-autoregressive decoder
pretrained 3D geometric priors
๐Ÿ’ผ Related Jobs
No related jobs found.
C
Chunan Yu
School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing 210094, China
Tianrun Chen
Tianrun Chen
Zhejiang University
Computer Vision3D ReconstructionComputational ImagingLarge Vision-Language Model
F
Fu Shen
Nanjing Institute of Agricultural Mechanization, Ministry of Agriculture and Rural Affairs, Nanjing 210014, China
C
Cheng Chen
KOKONI3D, Moxin (Huzhou) Technology Co., Ltd., China
Lanyun Zhu
Lanyun Zhu
NTU, CityUHK, SUTD, BUAA
Multimodal LearningComputer VisionResource-efficient LearningLarge Vision-Language Model
Yang Yang
Yang Yang
Nanjing University of Science and Technology
Data MiningMulti Modal LearningIncremental Learning