🤖 AI Summary
This study addresses the insufficiently validated effectiveness of large language models (LLMs) in diagnosing student error patterns within K-12 mathematics tutoring, challenging the prevalent assumption that inter-model consensus can substitute for independent validity testing. Employing an operationalized diagnostic codebook, this work systematically evaluates LLMs’ classification validity across five categories of mathematical failure modes using Cohen’s kappa and Cronbach’s alpha, while comparing cross-model and human–machine agreement. The findings reveal that high inter-model consensus actually obscures low validity, with agreement among models significantly exceeding human–machine alignment. These results robustly challenge the “consensus equals validity” assumption, establishing a methodological standard requiring independent human verification for evaluating LLMs in educational contexts.
📝 Abstract
In K-12 mathematics tutoring, student-tutor dialogue provides rich evidence of learners' problem-solving processes and sources of difficulty. Learning analytics research increasingly relies on large language models (LLMs) to extract such information from dialogue for a variety of downstream tasks, including knowledge tracing, behavioral modeling, and diagnosis of student reasoning errors. However, the validity of these model-generated interpretations remains insufficiently understood. In this exploratory study, we examine the validity of LLM classifications of five student failure modes in mathematics tutoring dialogue using an operational diagnostic codebook: uncertainty, misattribution, operator selection, conceptual gap, and procedural slip. Across models, human-LLM agreement was moderate (kappa = .524-.597), while cross-model agreement was substantially higher (kappa = .755-.781; alpha = .769). These findings show that cross-model agreement can create a misleading appearance of correctness, challenging the assumption that consensus among LLMs constitutes evidence of valid learner interpretation. For learning analytics, the implication is clear: scalable labeling is useful only if the inferred constructs are valid, and model consensus cannot substitute for independent evidence of that validity.