Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical limitation in current large language models (LLMs) employed as AI tutors: their frequent disregard for learners’ prior knowledge and instructional progression, coupled with a lack of effective evaluation of pedagogical suitability. To bridge this gap, the authors propose the Pedagogical Suitability Index (PSI), a theoretically grounded metric comprising six sub-dimensions that quantifies the alignment between LLM-generated tutoring responses and learners’ readiness as well as curricular pacing. For the first time, PSI is leveraged as a structured feedback signal to guide LLMs in refining their outputs. Experimental results demonstrate that 82.3% of 62 initially low-scoring cases showed significant improvement under PSI guidance, with human evaluators confirming the pedagogical validity of these enhancements—thereby transcending conventional evaluation paradigms that focus solely on answer correctness.
📝 Abstract
Large language models (LLMs) are increasingly used as AI tutors, but a correct answer is not always a pedagogically appropriate one. In classroom learning, effective help depends not only on correctness, but also on whether a response matches the learner's current foundation, the course sequence, and the timing of concept introduction. Existing evaluations focus mainly on answer quality, leaving this instructional fit under-measured. We present the Pedagogical Suitability Index (PSI), a composite metric of six theory-informed sub-scores that evaluates how well LLM-generated tutoring responses align with learner readiness and curricular progression, and we further use PSI as a structured feedback signal for response improvement. We evaluate four LLM tutors (ChatGPT, Gemini, Gemma4, and Qwen3) across 240 scenario-based evaluations using paired standard and defective prompts, then apply a PSI-guided regeneration protocol to 62 weak-performing cases. Baseline differences across the four tested models were modest overall (PSI range: 0.557 to 0.638), and open-weight and closed models did not exhibit a clear separation in pedagogical fit. Under the tested prompt perturbations, overall PSI remained largely stable (Delta = -0.002), though sub-score trade-offs emerged. More importantly, PSI-guided feedback substantially improved weak-performing cases: 51 of 62 cases improved (82.3%). Focused manual evaluation of the 62 PSI-selected weak cases provides initial evidence that the identified weaknesses are instructionally meaningful and that many PSI-guided regenerations correspond to human-judged improvement. These results suggest that learner- and curriculum-aware alignment may matter more for effective tutoring than model category alone, and that such alignment is both measurable and improvable.
Problem

Research questions and friction points this paper is trying to address.

pedagogical fit
AI tutors
large language models
instructional alignment
curricular progression
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pedagogical Suitability Index
LLM-based AI Tutoring
Instructional Alignment
Curriculum-aware Evaluation
Response Regeneration