๐ค AI Summary
This study evaluates the practical competency of large language models (LLMs) in providing corrective feedback and pedagogical explanations for English language teaching. Integrating retrieval-augmented generation with parameter sensitivity analysis, the research employs a multidimensional evaluation framework combining automated metrics and expert human assessment. Results indicate that while LLMs demonstrate robust surface-level error correction capabilities, they exhibit significant deficiencies in pedagogical terminology and domain-specific knowledge, revealing a critical dichotomy between strong correction skills and weak instructional competence. These findings identify key limitations of current LLMs in educational applications and provide empirical evidence to inform the optimization of intelligent tutoring systems, highlighting the necessity for enhanced domain adaptation in AI-driven language education tools.
๐ Abstract
While various organizations now actively encourage LLM use in classrooms, we still lack rigorous, systematic evaluations of how well these models actually perform the fundamental tasks of language pedagogy. This paper examines whether state-of-the-art LLMs can deliver the kind of corrective feedback and methodological explanations that language learners need. The study tests multiple large language models on their ability to identify, correct, and explain common learner mistakes in English, by systematically varying model parameters to investigate how these technical adjustments affect output quality, pedagogical clarity, and consistency, along with using retrieval-augmented generation to query methodological data. The evaluation employs automated metrics (GLEU, BERTScore) but also human expert judgments to capture dimensions that purely computational measures miss: linguistic nuance, cultural sensitivity, and instructional appropriateness. While models demonstrate impressive surface-level correction abilities, their explanations often lack the terminological and domain knowledge that effective language teaching requires, suggesting that current enthusiasm for AI-assisted language learning may be outpacing our understanding of these systems' actual pedagogical competence.