The learner who does not learn: when optimizing a pedagogical metric degrades LLM tutoring

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the counterproductive effects of fine-tuning large language model (LLM) tutoring systems on pedagogical adaptability metrics, which inadvertently leads to behavioral rigidity and degraded instructional quality. We propose the theoretical hypothesis that independent decision-making averaged over metrics tends to converge toward repetitive optimal solutions, and establish design principles for AI tutor benchmarks accordingly. An empirical investigation is conducted integrating automated scoring, open-source model fine-tuning, blinded expert evaluation, and weight analysis. Our findings reveal a critical over-optimization effect wherein improvements in metric scores are accompanied by significant declines in expert assessments, demonstrating that measurement validity does not entail optimization validity. This work provides essential caveats and actionable design guidelines for deploying LLMs in educational applications, cautioning against uncritical reliance on proxy metrics during alignment.
📝 Abstract
It is assumed that a natural way to improve the pedagogical quality of large language model tutors is to define a metric of instructional performance and fine-tune the model against it. To test this strategy, we designed a metric of pedagogical adaptivity that scores each instructional decision in a learning sequence against the conditions of the learning situation, which is the standard used for automated pedagogical scoring. We audited a frontier tutor across 2,000 learner scenarios, corrected its weakest cases by fine-tuning an open-weights proxy, and asked 31 trained educators to rate the pedagogical alignment of the outputs blind, before and after correction. The metric increased from +0.05 to +0.42 for the corrected cases, while the expert ratings decreased from 4.46 to 3.03, with the unmodified controls remaining unchanged and a base-proxy control ruling out the change of model. The tutor performed worse because any metric that scores decisions independently and averages them is maximized by repeating the single best decision, and the fine-tuned model collapsed to that exact optimum in every case, in and out of sample. Educators identified the repetition, which such metrics cannot represent, and preserving the learner's trajectory in the score reduced, but did not reverse, the metric's verdict. Weight analysis traced the correction to the model's output projection, where it had memorized its training strings rather than learned to adapt. We conclude that measurement validity does not imply optimization validity, and we derive design principles for benchmarks that assess or train AI-tutor instruction.
Problem

Research questions and friction points this paper is trying to address.

LLM tutoring
pedagogical metric
metric optimization
model collapse
measurement validity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pedagogical adaptivity metric
Goodhart's law in LLMs
Model collapse
Weight analysis
AI tutoring optimization