Evaluation Framework for AI Creativity: A Case Study Based on Story Generation

📅 2026-01-07
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing reference-based evaluation metrics struggle to capture the subjective and creative aspects of AI-generated stories. To address this limitation, this work proposes the first hierarchical, multidimensional framework for assessing creativity, encompassing four dimensions: novelty, value, adherence, and resonance. By employing Spike Prompting to control generated content and conducting a crowdsourced study with 115 participants, the research investigates how human judgments of creativity evolve across immediate and reflective evaluation phases. The findings reveal that reflective assessment significantly alters scoring outcomes and enhances inter-rater agreement. This framework effectively uncovers creativity dimensions overlooked by conventional metrics, substantially improving the granularity and reliability of creativity evaluation in AI-generated narratives.

Technology Category

Humans and AI: Game Design — Procedural Content Generation & StorytellingCognitive Modeling & Cognitive Systems: Computational CreativityNatural Language Processing: Generation

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Metrics for user behavior and evaluating successEconomics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAI
📝 Abstract
Evaluating creative text generation remains a challenge because existing reference-based metrics fail to capture the subjective nature of creativity. We propose a structured evaluation framework for AI story generation comprising four components (Novelty, Value, Adherence, and Resonance) and eleven sub-components. Using controlled story generation via ``Spike Prompting''and a crowdsourced study of 115 readers, we examine how different creative components shape both immediate and reflective human creativity judgments. Our findings show that creativity is evaluated hierarchically rather than cumulatively, with different dimensions becoming salient at different stages of judgment, and that reflective evaluation substantially alters both ratings and inter-rater agreement. Together, these results support the effectiveness of our framework in revealing dimensions of creativity that are obscured by reference-based evaluation.
Problem

Research questions and friction points this paper is trying to address.

AI creativity
story generation
evaluation metrics
subjective evaluation
creative text generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

creativity evaluation
story generation
structured framework
Spike Prompting
human judgment
🔎 Similar Papers
2021-04-06ACM Computing SurveysCitations: 32
💼 Related Jobs
No related jobs found.
P
Pharath Sathya
Graduate School of Informatics, Kyoto University, Kyoto, Japan
Y
Yin Jou Huang
Graduate School of Informatics, Kyoto University, Kyoto, Japan
F
Fei Cheng
Graduate School of Informatics, Kyoto University, Kyoto, Japan