Evaluating Interactivity: Toward Automated Assessment of AI-Generated Explorable Explanations

📅 2026-06-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing approaches struggle to effectively evaluate the quality of dynamic interactions in AI-generated explorable explanations, particularly lacking metrics for learner-controlled state transitions and context-sensitive feedback. This work proposes EE-Eval, a novel framework that formalizes interactivity as a finite state machine (FSM) and enables automated, multidimensional assessment by comparing the structural and semantic alignment between generated content and an ideal pedagogical FSM grounded in instructional intent. Integrating graph similarity with embedding-based semantic analysis, EE-Eval overcomes the limitations of conventional methods that focus narrowly on code executability or visual fidelity. Extensive experiments across 127 concepts and thousands of explanations generated by six AI models demonstrate that EE-Eval significantly outperforms existing baselines and exhibits strong agreement with human judgments of both interactivity and instructional effectiveness.
📝 Abstract
While large language models now enable rapid generation of interactive learning materials, evaluating the interaction quality of these explorable explanations remains an open challenge. Existing benchmarks largely focus on code executability or visual fidelity, providing limited insight into dynamic interaction behaviors such as learner-controlled state transitions and context-sensitive system responses, which are factors that critically shape learners' conceptual understanding. We present EE-Eval, an automated evaluation framework that formalizes interactivity as a finite space of learner-controllable states and transitions, represented as a Finite State Machine (FSM). By extracting FSMs from AI-generated explorable explanations, EE-Eval externalizes implicit interaction logic into an explicit, machine-interpretable graph. Evaluation is performed by comparing each generated FSM to an ideal FSM that encodes pedagogical intent, using a combination of graph-based metrics and embedding-based comparison of states, actions, and feedback to measure their structural and semantic similarity. Across thousands of generated explorable explanations spanning 127 concepts and produced by 6 AI models, EE-Eval consistently differentiates interaction quality beyond surface-level criteria such as functional correctness or visual quality, and exhibits substantially stronger alignment with human judgments of interactivity and pedagogical effectiveness than existing baselines. By framing interactivity as testable behavioral models rather than an emergent byproduct of LLM generation, EE-Eval transforms evaluation into a reflective diagnostic tool, enabling pedagogically grounded and actionable human-AI collaboration in creating interactive educational content.
Problem

Research questions and friction points this paper is trying to address.

interactivity
explorable explanations
AI-generated educational content
interaction quality
automated evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

explorable explanations
interactivity evaluation
Finite State Machine
automated assessment
pedagogical alignment
🔎 Similar Papers
No similar papers found.
X
Xiaozao Wang
New York University Shanghai
Z
Zhewei Wang
New York University Shanghai
H
Hongyi Wen
New York University Shanghai