LLM Prompt Evaluation for Educational Applications

📅 2026-01-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic, evidence-driven approaches for evaluating the effectiveness of large language model prompts in educational contexts, where balancing personalization and pedagogical alignment remains challenging. The authors propose a generalizable prompt evaluation framework that integrates six pedagogically informed prompt templates designed to generate follow-up questions within structured dialogues. For the first time in educational prompt engineering, they introduce tournament-style evaluation combined with the Glicko-2 rating system, complemented by multidimensional human assessments—covering format, conversational support, and learner adaptability—and validated through real user interaction data. Across 120 authentic interactions, a prompt template incorporating role specification, contextual management, and metacognitive strategies significantly outperformed others, achieving pairwise win rates of 81%–100%, thereby advancing prompt design from an intuition-based practice toward an evidence-driven paradigm.

Technology Category

Natural Language Processing: Prompt Engineering / PromptingMachine Learning: Large Multimodal Models (LMMs)Search and Optimization: Learning to Search

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
As large language models (LLMs) become increasingly common in educational applications, there is a growing need for evidence-based methods to design and evaluate LLM prompts that produce personalized and pedagogically aligned out-puts. This study presents a generalizable, systematic approach for evaluating prompts, demonstrated through an analysis of LLM-generated follow-up questions in a structured dialogue activity. Six prompt templates were designed and tested. The templates incorporated established prompt engineering patterns, with each prompt emphasizing distinct pedagogical strategies. The prompt templates were compared through a tournament-style evaluation framework that can be adapted for other educational applications. The tournament employed the Glicko2 rating system with eight judges evaluating question pairs across three dimensions: format, dialogue support, and appropriateness for learners. Data was sourced from 120 authentic user interactions across three distinct educational deployments. Results showed that a single prompt related to strategic reading out-performed other templates with win probabilities ranging from 81% to 100% in pairwise comparisons. This prompt combined persona and context manager pat-terns and was designed to support metacognitive learning strategies such as self-directed learning. The methodology showcases how educational technology re- searchers can systematically evaluate and improve prompt designs, moving beyond ad-hoc prompt engineering toward evidence-based prompt development for educational applications.
Problem

Research questions and friction points this paper is trying to address.

LLM prompt evaluation
educational applications
pedagogical alignment
personalized output
prompt design
Innovation

Methods, ideas, or system contributions that make the work stand out.

prompt evaluation
educational applications
tournament-style evaluation
Glicko2 rating system
metacognitive learning
🔎 Similar Papers
No similar papers found.
L
Langdon Holmes
Vanderbilt University, Nashville, Tennessee
A
Adam Coscia
Georgia Institute of Technology, Atlanta, Georgia
S
Scott A. Crossley
Vanderbilt University, Nashville, Tennessee
J
J. Choi
Vanderbilt University, Nashville, Tennessee
W
Wesley Morris
Vanderbilt University, Nashville, Tennessee