How Well Can LLMs Simulate Real Learner Evaluations of Educational Feedback?

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the capability boundaries of large language models (LLMs) in simulating real learners’ subjective evaluations of educational feedback. Leveraging high school biology feedback data, the research systematically assesses the simulation validity and statistical consistency of six LLMs at both individual and population levels by incorporating personalized learner profiles and few-shot prompt engineering. Results indicate that the overall simulation capacity of LLMs remains limited: while providing learner profiles improves score calibration at the individual level, it fails to enhance consistency at the population level. This work is the first to reveal the differential effects of distinct information adaptation strategies on LLMs’ ability to simulate learners’ subjective preferences, offering critical insights into the constraints of using generative AI for educational assessment tasks.
📝 Abstract
While recent studies have explored human behavior and preference simulation using large language models (LLMs), it remains unclear how well LLMs can simulate subjective evaluations from real learners in educational settings. We investigate this question using real learner evaluation data on feedback for high-school biology questions at both the group and individual levels. We compare performance with and without learner-specific information, such as personality traits and evaluation examples, across six models. Our results show that LLMs still have a limited ability to simulate learner evaluations. Providing learner profiles and examples improves score calibration and individual-level simulation, but more often fails to improve group-level consistency. These findings highlight the need to investigate which learner information and adaptation strategies are effective for learner preference simulation.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Learner Evaluation Simulation
Educational Feedback
Preference Simulation
Subjective Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Learner Evaluation Simulation
Educational Feedback
Learner Profiles
Preference Simulation
🔎 Similar Papers
No similar papers found.