AUV-Bench: Aesthetic Understanding and Generation Evaluation for User Interfaces

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing UI aesthetic assessments typically rely on isolated evaluations, making it difficult to verify the consistency between model judgments and actions. This work collaborates with professional designers to construct a benchmark encompassing four tasks—scoring, diagnosis, repair, and generation—and introduces a novel "judgment-action alignment" evaluation paradigm through a shared principle pool and controlled degradation instances. Based on 1,395 executable web interfaces and 660 degraded samples, this study systematically evaluates twelve multimodal large language models. The results reveal a significant judgment-action gap: while overall scoring performance is acceptable, precise diagnostic accuracy reaches only 24.7%, and correct judgments do not necessarily lead to successful repairs. These findings expose the current inability of such models to maintain coherent aesthetic understanding across evaluative and generative tasks.
📝 Abstract
Multimodal foundation models are increasingly used for evaluating and generating user interfaces (UIs), often producing seemingly reasonable aesthetic judgments and visually plausible pages. However, under professional design scrutiny, their behavior can differ substantially from that of human designers. In professional design practice, designers rely on a systematic set of aesthetic principles that consistently guide judgment, diagnosis, repair, and creation. A coherent aesthetic capability should therefore connect aesthetic judgment with design actions. Existing evaluations, however, typically assess these abilities in isolation, making it difficult to determine whether task-level success reflects a shared aesthetic understanding or merely fragmented task-specific competence. To address this gap, we introduce AUV-Bench, developed in collaboration with professional UI designers around 1,395 executable web interfaces and four tasks: aesthetic scoring, diagnosis, repair, and text-to-UI generation. The tasks share a pool of UIs and aesthetic principles, with diagnosis and repair further aligned on 660 controlled-degradation instances to enable instance-level analysis of judgment and action. Evaluation of 12 models reveals a capability imbalance: models show moderate agreement with professional designers in holistic aesthetic scoring, yet exact diagnosis-chain success peaks at only 24.7%. On the aligned diagnosis-repair cases, correct judgments and successful repairs do not consistently coincide, exposing a Judgment-Action Gap between identifying aesthetic problems and successfully acting on them. In open-ended generation, even leading models achieve only moderate aesthetic quality under human-calibrated evaluation. Overall, current models exhibit partial aesthetic competence, but still lack the fine-grained understanding and judgment-action coherence required for reliable UI design.
Problem

Research questions and friction points this paper is trying to address.

User Interface Aesthetics
Multimodal Foundation Models
Judgment-Action Gap
Aesthetic Evaluation
UI Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

AUV-Bench
Judgment-Action Gap
Multimodal Foundation Models
UI Aesthetic Evaluation
Controlled-Degradation
🔎 Similar Papers
No similar papers found.