🤖 AI Summary
This study investigates the feasibility of leveraging multimodal large language models (MLLMs) to automatically assess visual creativity in images without human annotations or model fine-tuning, while providing interpretable justifications. We present the first systematic evaluation of six MLLMs—Gemini, GPT, GLM, Kimi, Qwen, and Gemma—in a zero-shot setting for scoring the creativity of both AI-generated images and hand-drawn sketches. Experimental results demonstrate that model-generated scores exhibit significant correlation with human judgments, achieving a maximum Pearson correlation coefficient of 0.68. Furthermore, the models’ reasoning processes are interpretable, although this interpretability does not further improve scoring consistency with human assessments. Our findings highlight both the promising potential and inherent limitations of MLLMs for unsupervised evaluation of visual creativity.
📝 Abstract
Evaluating the originality of visual images poses enduring challenges for creativity assessment. Automated scoring using AI models has proven effective in the verbal domain, yet key questions remain about evaluating visual creativity and understanding how models arrive at their ratings. The present research asks whether multimodal large language models (LLMs) can serve as judges of visual creativity zero-shot (without any fine-tuning or examples of human ratings) and whether their "reasoning" output offers an interpretable window into their evaluation process. We tested six multimodal LLMs (Gemini 3 Flash, Gemma 4 31B IT, GPT-5.4 Mini, GLM-5v Turbo, Kimi K2.5, and Qwen 3.6 Plus) on 992 AI-generated images (based on human-written prompts) and 1,500 hand-drawn sketches scored for creativity by human raters. In Study 1, all models showed substantial alignment with human creativity ratings on both datasets (r = .57-.68 on AI-generated images; r = .29-68 on sketches). In Study 2, we analyzed the step-by-step reasoning processes of three LLMs evaluating the same images and drawings. Although reasoning made model evaluations interpretable -- showing what they attend to, how they balance originality vs. quality, and how they justify their ratings -- reasoning did not improve alignment with human ratings. In sum, our findings indicate that multimodal LLMs can match human judgments of visual creativity without any additional training, and that their reasoning reveals how AI models evaluate creativity. An open scoring app implementing this pipeline is available at https://review-visual-eval-scoring.hf.space.