🤖 AI Summary
This study addresses the lack of systematic, evidence-driven approaches for evaluating the effectiveness of large language model prompts in educational contexts, where balancing personalization and pedagogical alignment remains challenging. The authors propose a generalizable prompt evaluation framework that integrates six pedagogically informed prompt templates designed to generate follow-up questions within structured dialogues. For the first time in educational prompt engineering, they introduce tournament-style evaluation combined with the Glicko-2 rating system, complemented by multidimensional human assessments—covering format, conversational support, and learner adaptability—and validated through real user interaction data. Across 120 authentic interactions, a prompt template incorporating role specification, contextual management, and metacognitive strategies significantly outperformed others, achieving pairwise win rates of 81%–100%, thereby advancing prompt design from an intuition-based practice toward an evidence-driven paradigm.
📝 Abstract
As large language models (LLMs) become increasingly common in educational applications, there is a growing need for evidence-based methods to design and evaluate LLM prompts that produce personalized and pedagogically aligned out-puts. This study presents a generalizable, systematic approach for evaluating prompts, demonstrated through an analysis of LLM-generated follow-up questions in a structured dialogue activity. Six prompt templates were designed and tested. The templates incorporated established prompt engineering patterns, with each prompt emphasizing distinct pedagogical strategies. The prompt templates were compared through a tournament-style evaluation framework that can be adapted for other educational applications. The tournament employed the Glicko2 rating system with eight judges evaluating question pairs across three dimensions: format, dialogue support, and appropriateness for learners. Data was sourced from 120 authentic user interactions across three distinct educational deployments. Results showed that a single prompt related to strategic reading out-performed other templates with win probabilities ranging from 81% to 100% in pairwise comparisons. This prompt combined persona and context manager pat-terns and was designed to support metacognitive learning strategies such as self-directed learning. The methodology showcases how educational technology re- searchers can systematically evaluate and improve prompt designs, moving beyond ad-hoc prompt engineering toward evidence-based prompt development for educational applications.