VSA:Visual-Structural Alignment for UI-to-Code

📅 2025-12-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing design-to-code approaches generate flat, unstructured code lacking componentization support, resulting in low cohesion, high coupling, and poor maintainability. To address this, we propose a vision-structure alignment paradigm for end-to-end generation of modular frontend code from UI screenshots. Our method comprises three key innovations: (1) a spatially aware Transformer that explicitly models geometric relationships among interface elements; (2) heuristic UI pattern matching to identify reusable, semantically meaningful component structures; and (3) a schema-driven LLM synthesis engine that generates type-safe, framework-compliant (React/Angular) componentized code. Evaluated on multiple benchmarks, our approach significantly improves code modularity and architectural consistency, surpassing state-of-the-art methods. It is the first to reliably map pixel-level UI inputs to production-grade, maintainable frontend engineering artifacts.

Technology Category

Computer Vision: Multi-modal VisionNatural Language Processing: Code Generation / Program Synthesis from Natural LanguageHumans and AI: Intelligent User Interfaces

Application Category

Graph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsEconomics, Online Markets and Human Computation: Architectures and workflows that use LLMs for crowd workResponsible Web: Human-perceived consequences of algorithmic deployment on the web
📝 Abstract
The automation of user interface development has the potential to accelerate software delivery by mitigating intensive manual implementation. Despite the advancements in Large Multimodal Models for design-to-code translation, existing methodologies predominantly yield unstructured, flat codebases that lack compatibility with component-oriented libraries such as React or Angular. Such outputs typically exhibit low cohesion and high coupling, complicating long-term maintenance. In this paper, we propose extbf{VSA (VSA)}, a multi-stage paradigm designed to synthesize organized frontend assets through visual-structural alignment. Our approach first employs a spatial-aware transformer to reconstruct the visual input into a hierarchical tree representation. Moving beyond basic layout extraction, we integrate an algorithmic pattern-matching layer to identify recurring UI motifs and encapsulate them into modular templates. These templates are then processed via a schema-driven synthesis engine, ensuring the Large Language Model generates type-safe, prop-drilled components suitable for production environments. Experimental results indicate that our framework yields a substantial improvement in code modularity and architectural consistency over state-of-the-art benchmarks, effectively bridging the gap between raw pixels and scalable software engineering.
Problem

Research questions and friction points this paper is trying to address.

Generates modular UI code from visual designs
Improves code cohesion and reduces coupling
Aligns visual inputs with structural component templates
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical tree reconstruction via spatial-aware transformer
Modular templates from algorithmic pattern-matching of UI motifs
Schema-driven synthesis for type-safe, prop-drilled components
💼 Related Jobs
No related jobs found.
X
Xian Wu
Nanjing University
M
Ming Zhang
Nanjing University
Z
Zhiyu Fang
Nanjing University
F
Fei Li
Nanjing University
B
Bin Wang
Nanjing University
Y
Yong Jiang
Nanjing University
H
Hao Zhou
Nanjing University