Advancing vision-language models in front-end development via data synthesis

📅 2025-03-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Modern frontend frameworks (e.g., React, Vue) employ declarative rendering, state synchronization, and component reuse—introducing substantial complexity in generating executable code from design mockups and limiting the performance of existing vision-language models (VLMs). Method: We propose a reflective agent workflow and introduce three novel image-code paired data synthesis strategies: evolutionary (enhancing diversity), waterfall-model (ensuring logical consistency), and incremental development (progressively increasing complexity). We establish a “vision-first” VLM training paradigm—“observe image before generating code”—and integrate self-contained code extraction, multi-stage rendering, and descriptive caption generation. Our approach is instantiated and evaluated on the Flame VLM. Contribution/Results: Experiments demonstrate significant improvements over baselines in React code generation. Pass@k evaluation confirms that vision-prior understanding substantially enhances both functional correctness and overall code quality.

Technology Category

Computer Vision: Large Vision ModelsNatural Language Processing: Code Generation / Program Synthesis from Natural LanguageMachine Learning: Large Multimodal Models (LMMs)

Application Category

Graph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Modern front-end (FE) development, especially when leveraging the unique features of frameworks like React and Vue, presents distinctive challenges. These include managing modular architectures, ensuring synchronization between data and visual outputs for declarative rendering, and adapting reusable components to various scenarios. Such complexities make it particularly difficult for state-of-the-art large vision-language models (VLMs) to generate accurate and functional code directly from design images. To address these challenges, we propose a reflective agentic workflow that synthesizes high-quality image-text data to capture the diverse characteristics of FE development. This workflow automates the extraction of self-containedfootnote{A extbf{self-contained} code snippet is one that encapsulates all necessary logic, styling, and dependencies, ensuring it functions independently without requiring external imports or context.} code snippets from real-world projects, renders the corresponding visual outputs, and generates detailed descriptions that link design elements to functional code. To further expand the scope and utility of the synthesis, we introduce three data synthesis strategies: Evolution-based synthesis, which enables scalable and diverse dataset expansion; Waterfall-Model-based synthesis, which generates logically coherent code derived from system requirements; and Additive Development synthesis, which iteratively increases the complexity of human-authored components. We build a large vision-language model, Flame, trained on the synthesized datasets and demonstrate its effectiveness in generating React code via the $ ext{pass}@k$ metric. Our results suggest that a code VLM trained to interpret images before code generation may achieve better performance.
Problem

Research questions and friction points this paper is trying to address.

Challenges in generating accurate code from design images using VLMs.
Difficulties in managing modular architectures and data-visual synchronization.
Need for high-quality image-text data synthesis for FE development.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reflective agentic workflow synthesizes image-text data
Automates extraction of self-contained code snippets
Introduces three data synthesis strategies for scalability
T
Tong Ge
KE Holdings Inc.
Y
Yashu Liu
KE Holdings Inc.
J
Jieping Ye
Alibaba Group
T
Tianyi Li
KE Holdings Inc.
C
Chao Wang
KE Holdings Inc.