🤖 AI Summary
This work addresses the high cost, lengthy timelines, and limited fidelity of traditional UI/UX evaluation methods—such as user studies and A/B testing—in simulating authentic user feedback. To overcome these limitations, the authors propose PerceptUI, a novel framework that enables fine-grained, persona-based UI/UX feedback generation for the first time. Leveraging multimodal large language models, PerceptUI integrates persona-conditioned prompting, contrastive reflection-based fine-tuning, and a failure-trajectory-driven prompt evolution mechanism to distill rational justifications from human decision-making processes, thereby enhancing the model’s introspective capabilities. Experimental results demonstrate that PerceptUI achieves human-level authenticity in generated feedback across multiple domains and datasets, generalizes effectively to unseen interface issues and user personas, and supports the synthesis of population-level response distributions.
📝 Abstract
User interface (UI) and user experience (UX) evaluation is central to product development, yet reliable feedback still relies on recruiting human participants or running online A/B tests, making early-stage iteration slow and costly. In light of this, recent work has explored Multimodal Large Language Models as proxy evaluators. However, existing approaches either produce surface-level critiques or a judgment that reflects the model's own biases rather than the genuine response of a particular user. We introduce PerceptUI, a framework for persona-conditioned UI/UX evaluation that predicts how a specific user would answer interface-related questions and produces natural-language rationales. PerceptUI is trained in two stages: (i) contrastive reflection fine-tuning distills teacher-generated rationales by extracting lessons from human decisions, and (ii) a reflective prompt-evolution step from the model's own failure traces. Across multiple domains and datasets, PerceptUI achieves human-level realism, generalizes to unseen questions and personas, and yields population-level response distributions.