🤖 AI Summary
This work addresses the challenge of preserving geometric fidelity and cross-view spatial alignment in engineering drawings with existing models by proposing a large language model-driven framework for multi-view vector drawing generation. The framework unifies the generation task as sequence modeling, eliminating the need for raster encoders. It introduces a novel hierarchical suffix tokenization representation alongside a progressive curriculum scheduling strategy, enabling a smooth transition from local structural inpainting to macro-level generation. Experimental results demonstrate that the proposed method significantly improves both the geometric fidelity and syntactic accuracy of generated drawings across various conditional and unconditional generation tasks.
📝 Abstract
Scalable Vector Graphics (SVG) are essential for modern industrial Computer-Aided Design (CAD). However, existing autoregressive SVG generation models are predominantly tailored for artistic creation and struggle to maintain the rigorous geometric fidelity and cross-view spatial alignment required for engineering drawings. To bridge this gap, we introduce \textbf{DrawingsDreamer}, a unified Large Language Model (LLM)-driven framework for multi-view vector-based engineering drawings generation. By formulating the generation of multi-view engineering drawings purely as a sequence modeling task, we eliminate the need of raster image encoders. We propose a Streamlined Representation utilizing hierarchical postfix tokenization, which guides the model to establish local geometric coordinates before assigning semantic boundaries. Optimized via a progressive task-aware curriculum schedule, \textbf{DrawingsDreamer} effectively transitions from localized structural repair to macroscopic generation in a unified model. Extensive experiments demonstrate that our unified model achieves strong performance in both geometric fidelity and syntactic accuracy across diverse conditional and unconditional generation tasks.