iterative visual refinement

Designs and implements closed-loop pipelines that render and analyze multi-view visualizations of procedural executions or geometric models to detect and localize geometric and procedural mismatches. Uses those visual diagnostics to iteratively revise CAD representations, intermediate representations, or generation procedures and re-execute the workflow until the produced boundary representation or visualization meets the required fidelity or correctness.

iterativevisualrefinement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Formalizing Linear Motion G-code for Invariant Checking and Differential Testing of Fabrication Tools

Aug 31, 2025
YH
Yumeng He
🏛️ University of Utah | Certora Inc. | University of Washington | University of Rochester

The absence of formal verification methods for G-code linear motion in 3D printing hinders assurance of geometric-code consistency. Method: This paper proposes a dimension-elevated semantic representation framework that parses G-code into sets of axis-aligned bounding boxes and their approximated point clouds, integrating geometric modeling with program analysis to enable invariant checking and differential testing across the manufacturing pipeline. Contribution/Results: The framework supports, for the first time, quantitative cross-slicer comparison (Cura vs. PrusaSlicer), error localization, and root-cause analysis of defects introduced during mesh repair (e.g., in MeshLab or Meshmixer). Evaluated on 58 real-world models, it efficiently detects slicing anomalies induced by small geometric features, exposes behavioral discrepancies among mainstream slicers, and identifies new errors inadvertently introduced during repair—thereby significantly enhancing the verifiability and reliability of end-to-end additive manufacturing pipelines.

Enabling error localization in CAD models through differential testingFormalizing G-code for invariant checking in fabrication toolsQuantitative comparison of slicers and mesh repair tool efficacy

This work addresses the lack of precise geometric validation in existing text-to-CAD generation methods, which often fail to correct dimensional inaccuracies. We propose a multi-agent collaborative framework that integrates inner-loop iterative correction of code execution errors with an outer-loop refinement combining exact geometric metrics from the OpenCASCADE kernel and global shape assessment via a vision-language model. Our approach uniquely unifies procedural geometric verification with visual-semantic judgment, yielding a retrieval-augmented generation (RAG) system that requires no fine-tuning and naturally evolves with CAD libraries. Evaluated on a newly curated benchmark of 100 multi-difficulty examples, our method achieves a 100% execution success rate, improves median IoU from 0.8085 to 0.9629, and reduces average Chamfer Distance from 28.37 to 0.74.

CAD model reliabilitydimensional errorsexecution errors

CADReview: Automatically Reviewing CAD Programs with Error Detection and Correction

May 28, 2025
JC
Jiali Chen
🏛️ South China University of Technology | The Hong Kong Polytechnic University

To address the low efficiency of manual CAD design review and the challenge of ensuring consistency between 3D models and reference images, this paper proposes ReCAD—the first end-to-end framework for automated CAD program review and repair. Methodologically, we construct CADReview, a large-scale paired dataset of CAD programs and corresponding images (20K+ samples), and design a customized architecture integrating multimodal understanding, program semantic parsing, and geometric spatial reasoning. We further introduce structured prompting and stepwise error correction to enable fine-grained error localization and program-level repair. Evaluated on the CADReview benchmark, ReCAD achieves a 32.7% improvement in error detection accuracy and a 28.4% increase in correction success rate over state-of-the-art multimodal large language models (MLLMs), demonstrating its effectiveness and practical potential for industrial design automation.

Automatically detect and correct errors in CAD programsEnsure consistency between 3D objects and reference imagesImprove accuracy of geometric component recognition in MLLMs

Generating CAD Code with Vision-Language Models for 3D Designs

Oct 07, 2024
KA
Kamel Alrashedy
🏛️ Georgia Institute of Technology

Large language models (LLMs) often generate CAD scripts that deviate from design intent due to geometric distortions or compilation failures. To address this, we propose CADCodeVerify—the first iterative vision-language verification framework tailored for CAD code generation. It leverages multimodal large models (e.g., GPT-4V) to automatically formulate verification-oriented visual questions, integrates CAD rendering, point-cloud Chamfer Distance (CD) evaluation, and feedback-driven multi-turn prompt optimization, enabling end-to-end refinement from natural language to high-fidelity, executable CAD code. Our contributions are threefold: (1) the first vision-closed-loop verification paradigm for CAD code generation; (2) CADPrompt—the first benchmark comprising 200 image-text–code triplets; and (3) on GPT-4, a 7.30% reduction in point-cloud CD distance and a 5.0% improvement in compilation success rate, significantly enhancing structural integrity and dimensional compliance.

Enhance 3D object structure and program success rate.Generate and validate CAD code using Vision-Language Models.Verify and improve 3D objects from CAD code.

Latest Papers

What's happening recently
View more

This work addresses the degradation in generation quality observed in large language model (LLM)-driven CAD agents during iterative repair, primarily caused by the loss of user requirements, operation history, and failure evidence. To mitigate this, the authors propose a persistent tracking mechanism that jointly models user intent, modeling steps, failure evidence, and candidate outputs. The approach enables reliable repair through diagnosing faulty operations, performing bounded edit searches within their local dependency regions, and integrating execution validation with retention checks. For the first time, the method introduces persistent state tracking, localized dependency-aware editing, and reusable skill memory, significantly improving both repair success rates and geometric fidelity. Evaluated on DeepCAD, the system achieves state-of-the-art performance: ablating persistent state reduces repair success by nearly 50%, while removing local search doubles geometric error and API call count; pre-populating the skill library effectively lowers retry frequency, token consumption, and latency.

CAD generationfault diagnosisgeometric quality

This study addresses the challenges of recovering sparse control cages from dense meshes, including inferring control requirements, missing feature curves, and difficulty in correcting initial results. To overcome these issues, this work proposes a modeler-inspired agent workflow that iteratively optimizes through planning and diagnosis phases. A stateful feedback mechanism is introduced to support repair or rollback decisions by incorporating semantic judgments into the pipeline. To ensure inspectability, the planner avoids directly generating vertices, instead integrating multi-view geometric evidence analysis with tool execution verification. Evaluated on a fixed benchmark queue, the proposed method outperforms automatic remeshing approaches across five metrics, reducing Chamfer-L1 to 0.443% and improving the F-score to 92.78%, thereby achieving highly controllable mesh-to-subdivision surface reconstruction.

Control cage recoveryMesh-to-SubD reconstructionReverse engineering

This study addresses the limitations of recovering executable CAD programs from 3D meshes, specifically the difficulty in handling diverse modeling operations and correcting geometric errors. To this end, we propose a generative optimization framework that introduces a pioneering step-wise generation mechanism. Our method leverages large language models to predict modeling actions, integrating intermediate geometric states with an IoU-guided tree search to enable local editing and iterative refinement. Furthermore, we release ARCADE-1.5M, a large-scale dataset encompassing multi-operation CAD sequences. Extensive experiments demonstrate that our approach achieves state-of-the-art reconstruction accuracy across multiple benchmarks, yielding a relative IoU improvement of up to 87.2% on complex shapes while maintaining high program validity.

3D mesh to CADCAD program generationgeometric reconstruction

Existing CAD generation methods struggle to simultaneously preserve modeling history, topological reference stability, and feature-level editability in cross-platform scenarios. This work proposes CADIR—an agent-oriented, executable intermediate representation that explicitly constructs a procedural graph encompassing operation sequences, parameter dependencies, constraints, and topological selections based on the OpenCASCADE (OCCT) geometric kernel. To enable faithful cross-platform model reconstruction, CADIR introduces a geometric signature matching mechanism. It is the first approach to support explicit procedural graph representations that allow editing across heterogeneous CAD backends. By integrating text- or image-driven procedural graph retrieval, CADIR demonstrates high-fidelity, editable reuse of complete models and substructures across FreeCAD, SolidWorks, and Fusion 360, enabling seamless subsequent modifications.

CAD generationconstruction historycross-backend editing

Hot Scholars

ML

Ming Li

Professor, Duke Kunshan University
Speech ProcessingAudio ProcessingAffective ComputingBehavior Signal Processing
ET

Enzo Tartaglione

Associate Professor, Télécom Paris, Institut Polytechnique de Paris
deep learningcompressionpruningdebiasing
JR

Ji-Rong Wen

Gaoling School of Artificial Intelligence, Renmin University of China
Large Language ModelWeb SearchInformation RetrievalMachine Learning
YL

Yijun Li

Adobe Research
Computer Vision
PT

Philip Torr

Professor, University of Oxford
Department of Engineering