Score
Designs and implements methods to detect and assign semantic labels to geometric features in CAD models, such as identifying faces, edges, holes, and their topological relationships. Builds pipelines that parse boundary representations into face‑adjacency/topology graphs and apply feature‑recognition and graph‑based models (e.g., GraphSAGE) to inject semantic labels and produce geometry abstractions.
Industrial CAD models often lack semantic, spatial, and functional information, limiting their utility in robotic simulation and high-level scene understanding. This work proposes an offline method that leverages large vision-language models (LVLMs)—introduced for the first time into CAD environments—to automatically generate structured 3D scene graphs that explicitly model manipulable objects and their functional relationships. By effectively integrating semantic parsing with functional reasoning, the approach achieves high-precision semantic annotation and relationship recognition on industrial structures such as piping systems. Both qualitative and quantitative evaluations demonstrate its effectiveness. The associated code and dataset have been made publicly available.
Existing CAD learning approaches discretize B-Rep models into triangle meshes, thereby discarding the analytical surface representations and topological information essential for consistent instance-level analysis. This work proposes STEP-Parts, a deterministic pipeline that directly extracts geometric instance partitions from native STEP B-Rep data. The method defines partitions based on intrinsic B-Rep topology, merges faces using analytical surface types and near-tangent plane continuity criteria, and transfers labels to triangulated meshes via face-to-mesh correspondence mapping. STEP-Parts ensures boundary consistency across varying triangulations and processes the DeepCAD subset of the ABC dataset—comprising approximately 180,000 models—in under six hours. The resulting labels significantly enhance performance in implicit reconstruction-segmentation tasks and point cloud networks. Code and precomputed labels are publicly released.
High annotation cost and poor generalizability plague geometric feature labeling in engineering design images. Method: We propose a large-model-driven automated annotation framework featuring (1) GeoBiked—the first bicycle-structure-specific dataset (4,355 images); (2) Diffusion-Hyperfeatures, a novel representation technique enabling precise cross-image correspondence of geometric keypoints; (3) a dual-path GPT-4o input mechanism integrating image and category labels, enhanced by systematic prompt engineering and a structured annotation schema to ensure fidelity and consistency of technical descriptions; and (4) a multi-source image collaborative geometric localization strategy. Results: Experiments demonstrate significant improvements in keypoint detection accuracy on unseen samples, with high descriptive accuracy from GPT-4o outputs. The framework validates the feasibility and strong generalization capability of large language-vision models for engineering image understanding and fine-grained geometric annotation.
Vector polygon geometry classification faces challenges including inadequate discrete representation modeling, weak transformation robustness, and poor cross-domain generalization. This paper proposes PolyMP—the first permutation-invariant graph neural network tailored for vector polygons—explicitly modeling each polygon as a geometry-aware graph (vertices as nodes, edges as edges). Leveraging equivariant coordinate normalization and structured message passing, PolyMP achieves strong robustness to translation, rotation, scaling, shearing, and vertex deletion. Evaluated on synthetic glyph and real-world building contour datasets, PolyMP significantly outperforms rasterization-based baselines: it improves classification accuracy by over 8%, attains 95.2% robustness under vertex perturbations, and achieves 89.3% cross-domain transfer accuracy. To our knowledge, this is the first work to empirically validate the effectiveness and generalization capability of pure-vector graph neural networks for geometric shape classification.
This survey addresses the growing need for systematic understanding of geometric deep learning (GDL) in computer-aided design (CAD). We focus on three core challenges: design similarity analysis, 2D/3D model synthesis, and CAD reconstruction from point clouds or single/multi-view inputs. Methodologically, we propose the first task taxonomy tailored to CAD-specific GDL, rigorously delineating paradigm boundaries; consolidate over 12 benchmark datasets and 50+ open-source implementations; and unify diverse techniques—including graph neural networks, point cloud processing, mesh learning, generative modeling, and multimodal representation—into a coherent methodological framework. Our analysis identifies critical bottlenecks and open problems, while providing an extensible research roadmap and practical implementation guidelines. The work significantly advances the efficiency, generalizability, and real-world applicability of intelligent CAD design systems.
Traditional engineering design relies heavily on costly simulations, while existing data-driven approaches often overlook the parametric nature and design semantics of CAD models, limiting their integration into design workflows and interpretability. This work proposes Attribute Feature Graphs (AFGs), which, for the first time, encode native CAD features—such as extrusions and ribs—as graph nodes, with directed edges representing geometric and dependency relationships. This representation preserves design intent while enabling end-to-end learning with graph neural networks (GNNs). Evaluated on the CarHoods10K dataset, the resulting GNN surrogate model achieves prediction accuracy comparable to state-of-the-art methods and supports direct feature editing within CAD environments with real-time performance feedback. Crucially, the approach provides traceability and interpretability by linking predictions back to specific design features.
This work addresses the limitations of existing approaches to machining feature recognition in B-Rep CAD models—namely, low sample efficiency, high computational cost, and the decoupled evaluation of instance segmentation and semantic labeling—by introducing FeatureFox, a lightweight and efficient panoptic recognition pipeline. FeatureFox employs an enhanced binary edge classifier with enriched edge attributes to localize feature boundaries, extracts connected components from a pruned face adjacency graph to recover instances, and predicts machining categories via subgraph aggregation. Notably, it is the first to adopt the Panoptic Quality (PQ) metric for this task, enabling end-to-end joint optimization. With only around 250 training parts and seconds of training time, FeatureFox achieves PQ > 0.9, matching the performance of the deep learning model AAGNet on the full dataset while demonstrating superior feature localization accuracy and generalization capability.
Existing methods for CAD symbol recognition often overlook the deep semantic content of textual annotations, thereby limiting holistic detection performance. To address this limitation, this work proposes TextCAD, a multimodal framework that jointly models graphical primitives and textual annotations through a Type-Attribute Correlation Encoder (TACE) and a multi-level semantic alignment mechanism. This approach effectively integrates their compositional semantics and hierarchical structure. By incorporating primitive downsampling and cross-modal semantic injection strategies, the proposed method achieves state-of-the-art symbol recognition accuracy on real-world architectural drawing datasets, significantly outperforming existing approaches.
This work addresses the challenge faced by visually impaired users in accessing node-link diagrams commonly distributed as bitmap images, a task for which existing assistive technologies are ill-suited due to their reliance on structured data rather than visual input. The paper presents the first lightweight deep learning approach for semantic segmentation of such diagram images, training a compact model on a large-scale synthetic dataset to achieve pixel-level parsing. The proposed method attains over 93% pixel accuracy on synthetic data and demonstrates strong performance both quantitatively and qualitatively. By enabling precise extraction of diagram semantics directly from rasterized images, this approach establishes a viable foundation for non-visual interaction and effectively bridges a critical gap in accessibility technology for bitmap-based graphical content.
Existing CAD generation methods struggle to simultaneously preserve modeling history, topological reference stability, and feature-level editability in cross-platform scenarios. This work proposes CADIR—an agent-oriented, executable intermediate representation that explicitly constructs a procedural graph encompassing operation sequences, parameter dependencies, constraints, and topological selections based on the OpenCASCADE (OCCT) geometric kernel. To enable faithful cross-platform model reconstruction, CADIR introduces a geometric signature matching mechanism. It is the first approach to support explicit procedural graph representations that allow editing across heterogeneous CAD backends. By integrating text- or image-driven procedural graph retrieval, CADIR demonstrates high-fidelity, editable reuse of complete models and substructures across FreeCAD, SolidWorks, and Fusion 360, enabling seamless subsequent modifications.