🤖 AI Summary
This study addresses the bottleneck in CAD-to-assembly planning, which traditionally relies on manual intervention and supplementary metadata, by proposing an end-to-end automated framework that takes mesh models as its sole input. Methodologically, the approach generates optimal assembly sequences through physics-based disassembly simulation and Design for Assembly (DfA) cost function ranking, eliminating the need for joint or fastener annotations. Furthermore, a multimodal large language model is incorporated to automatically synthesize assembly manuals and recommend appropriate tools. Experimental results demonstrate that the proposed method reduces simulated assembly time by 35% and achieves a tool selection accuracy of 88.6%, significantly enhancing both the efficiency and intelligence of manufacturing planning processes.
📝 Abstract
Turning a CAD design into an assembly plan is still largely done by hand, requiring engineers to reason about geometric feasibility, tool access, stability, and the ergonomics of human assembly. In this work, we encode long-established design for assembly (DfA) principles into a contained, end-to-end approach for generating assembly plans. Our approach takes only a mesh assembly and produces either a step-by-step assembly manual or a structured failure report, requiring no joint metadata, fastener annotations, or additional information. Four major components of a manufacturing plan are addressed autonomously: an assembly tool list, the assembly sequence and subassemblies, an assembly manual, and design feedback for improving assemblability. For determining the sequence plan, we systematically disassemble the object in a physics simulator and apply a cost function that encodes DfA principles. Manual generation, tool labelling, and assembly feedback rely primarily on multimodal large language models. Compared with a baseline that always removes the outermost part first from Tian et al., DfA-aware sequence planning reduces simulated assembly time, measured with a robot-arm assembly-time proxy, by 35% on 136 assemblies of 5 to 30 parts. The correct tool is selected for 88.6% of assembly steps. A vision-language model judge compares the generated manuals against ablated variants, identifying which page elements carry the information a reader needs. The presented approach and open-source code are available for use by engineers or AI agents looking to rapidly accelerate the creation of manufacturing plans for a given product design.