π€ AI Summary
This study addresses the lack of quantifiable definitions and evaluation criteria for βcreativityβ in generative AI. We propose, for the first time, a behaviorally grounded framework for assessing practical creativity in image generation models. Methodologically, we introduce an interpretable, three-dimensional metric encompassing diversity, novelty, and appropriateness; integrate multi-model comparative experiments, human perceptual evaluation, and statistical significance testing to ensure reproducibility, comparability, and alignment with human intuition. Validation across mainstream image-to-image translation models demonstrates strong agreement between our metric rankings and human subjective scores (Spearman Ο > 0.85), significantly outperforming existing black-box evaluation approaches. The framework enables objective, cross-model comparison of creative performance and provides empirical guidance for users selecting optimal generative models according to task-specific requirements.
π Abstract
Creativity of generative AI models has been a subject of scientific debate in the last years, without a conclusive answer. In this paper, we study creativity from a practical perspective and introduce quantitative measures that help the user to choose a suitable AI model for a given task. We evaluated our measures on a number of popular image-to-image generation models, and the results of this suggest that our measures conform to human intuition.