🤖 AI Summary
This study addresses the challenge of constraint-aware reasoning in multimodal large language models (MLLMs) when integrating visual and metadata inputs for industrial-grade UI code generation. We construct a benchmark comprising 2,861 production-level designs and formally define three tasks: end-to-end generation, requirement inference, and implementation. Methodologically, we combine expert annotations with domain-specific rules to guide cross-modal reasoning. Evaluations across eight mainstream MLLMs reveal that visual reconstruction proficiency does not equate to code compliance, and current models remain inadequate for industrial deployment requirements. This work bridges the evaluation gap in this scenario, and the dataset has been open-sourced to facilitate future research.
📝 Abstract
A key challenge for multimodal large language models (MLLMs) is moving beyond visual recognition to constraint-aware cross-modal reasoning. This involves combining visual cues with information from other modalities to understand elements' relationships under domain-specific rules. This challenge is acutely evident in industrial design-to-code (D2C), which converts user interface (UI) designs into code and requires MLLMs to connect design images with disorganized layer metadata, infer component and layout implementation requirements, and realize them in code under target-library constraints. However, these capabilities remain insufficiently evaluated in realistic industrial settings. To fill this gap, we present TaoD2C-Bench, a benchmark for evaluating MLLMs' ability to generate UI code that satisfies implementation requirements in industrial applications. The TaoD2C dataset consists of 2,861 production designs from 17 commercial platforms with 97,652 expert annotations across four categories: Component, Group, Alignment, and Position. These annotations distinguish required constraints from permitted implementation choices. TaoD2C-Bench defines three tasks: end-to-end UI code generation, requirement inference, and requirement realization. Evaluating eight MLLMs reveals substantial gaps in generating UI code that satisfies implementation requirements, alongside distinct performance profiles in inference and realization. We further show that MLLMs' visual reconstruction ability does not necessarily imply an ability to generate code that meets these requirements. We release TaoD2C to support research on industrial UI code generation.