🤖 AI Summary
Existing design-to-code approaches generate flat, unstructured code lacking componentization support, resulting in low cohesion, high coupling, and poor maintainability. To address this, we propose a vision-structure alignment paradigm for end-to-end generation of modular frontend code from UI screenshots. Our method comprises three key innovations: (1) a spatially aware Transformer that explicitly models geometric relationships among interface elements; (2) heuristic UI pattern matching to identify reusable, semantically meaningful component structures; and (3) a schema-driven LLM synthesis engine that generates type-safe, framework-compliant (React/Angular) componentized code. Evaluated on multiple benchmarks, our approach significantly improves code modularity and architectural consistency, surpassing state-of-the-art methods. It is the first to reliably map pixel-level UI inputs to production-grade, maintainable frontend engineering artifacts.
📝 Abstract
The automation of user interface development has the potential to accelerate software delivery by mitigating intensive manual implementation. Despite the advancements in Large Multimodal Models for design-to-code translation, existing methodologies predominantly yield unstructured, flat codebases that lack compatibility with component-oriented libraries such as React or Angular. Such outputs typically exhibit low cohesion and high coupling, complicating long-term maintenance. In this paper, we propose extbf{VSA (VSA)}, a multi-stage paradigm designed to synthesize organized frontend assets through visual-structural alignment. Our approach first employs a spatial-aware transformer to reconstruct the visual input into a hierarchical tree representation. Moving beyond basic layout extraction, we integrate an algorithmic pattern-matching layer to identify recurring UI motifs and encapsulate them into modular templates. These templates are then processed via a schema-driven synthesis engine, ensuring the Large Language Model generates type-safe, prop-drilled components suitable for production environments. Experimental results indicate that our framework yields a substantial improvement in code modularity and architectural consistency over state-of-the-art benchmarks, effectively bridging the gap between raw pixels and scalable software engineering.