🤖 AI Summary
This study addresses a critical limitation in current large language models (LLMs) employed as AI tutors: their frequent disregard for learners’ prior knowledge and instructional progression, coupled with a lack of effective evaluation of pedagogical suitability. To bridge this gap, the authors propose the Pedagogical Suitability Index (PSI), a theoretically grounded metric comprising six sub-dimensions that quantifies the alignment between LLM-generated tutoring responses and learners’ readiness as well as curricular pacing. For the first time, PSI is leveraged as a structured feedback signal to guide LLMs in refining their outputs. Experimental results demonstrate that 82.3% of 62 initially low-scoring cases showed significant improvement under PSI guidance, with human evaluators confirming the pedagogical validity of these enhancements—thereby transcending conventional evaluation paradigms that focus solely on answer correctness.
📝 Abstract
Large language models (LLMs) are increasingly used as AI tutors, but a correct answer is not always a pedagogically appropriate one. In classroom learning, effective help depends not only on correctness, but also on whether a response matches the learner's current foundation, the course sequence, and the timing of concept introduction. Existing evaluations focus mainly on answer quality, leaving this instructional fit under-measured. We present the Pedagogical Suitability Index (PSI), a composite metric of six theory-informed sub-scores that evaluates how well LLM-generated tutoring responses align with learner readiness and curricular progression, and we further use PSI as a structured feedback signal for response improvement. We evaluate four LLM tutors (ChatGPT, Gemini, Gemma4, and Qwen3) across 240 scenario-based evaluations using paired standard and defective prompts, then apply a PSI-guided regeneration protocol to 62 weak-performing cases. Baseline differences across the four tested models were modest overall (PSI range: 0.557 to 0.638), and open-weight and closed models did not exhibit a clear separation in pedagogical fit. Under the tested prompt perturbations, overall PSI remained largely stable (Delta = -0.002), though sub-score trade-offs emerged. More importantly, PSI-guided feedback substantially improved weak-performing cases: 51 of 62 cases improved (82.3%). Focused manual evaluation of the 62 PSI-selected weak cases provides initial evidence that the identified weaknesses are instructionally meaningful and that many PSI-guided regenerations correspond to human-judged improvement. These results suggest that learner- and curriculum-aware alignment may matter more for effective tutoring than model category alone, and that such alignment is both measurable and improvable.