Agreement Is Not Validity: Cross-Model LLM Consensus in Diagnosing Student Failure Modes in K-12 Math Tutoring Dialogue

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficiently validated effectiveness of large language models (LLMs) in diagnosing student error patterns within K-12 mathematics tutoring, challenging the prevalent assumption that inter-model consensus can substitute for independent validity testing. Employing an operationalized diagnostic codebook, this work systematically evaluates LLMs’ classification validity across five categories of mathematical failure modes using Cohen’s kappa and Cronbach’s alpha, while comparing cross-model and human–machine agreement. The findings reveal that high inter-model consensus actually obscures low validity, with agreement among models significantly exceeding human–machine alignment. These results robustly challenge the “consensus equals validity” assumption, establishing a methodological standard requiring independent human verification for evaluating LLMs in educational contexts.
📝 Abstract
In K-12 mathematics tutoring, student-tutor dialogue provides rich evidence of learners' problem-solving processes and sources of difficulty. Learning analytics research increasingly relies on large language models (LLMs) to extract such information from dialogue for a variety of downstream tasks, including knowledge tracing, behavioral modeling, and diagnosis of student reasoning errors. However, the validity of these model-generated interpretations remains insufficiently understood. In this exploratory study, we examine the validity of LLM classifications of five student failure modes in mathematics tutoring dialogue using an operational diagnostic codebook: uncertainty, misattribution, operator selection, conceptual gap, and procedural slip. Across models, human-LLM agreement was moderate (kappa = .524-.597), while cross-model agreement was substantially higher (kappa = .755-.781; alpha = .769). These findings show that cross-model agreement can create a misleading appearance of correctness, challenging the assumption that consensus among LLMs constitutes evidence of valid learner interpretation. For learning analytics, the implication is clear: scalable labeling is useful only if the inferred constructs are valid, and model consensus cannot substitute for independent evidence of that validity.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Learning Analytics
Validity
Cross-Model Agreement
Math Tutoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-model consensus
Large language models
Learning analytics
Failure mode diagnosis
Validity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Clayton Cohn
Clayton Cohn
PhD Student, Vanderbilt University
NLPLLMsAIEDAgents
J
Joyce Fonteles
College of Connected Computing, Vanderbilt University, Nashville, USA.
K
Kirk Vanacore
College of Computing and Information Science, Cornell University, Ithaca, USA.
G
Gianni Mazza
Third Space Learning, Swindon, UK.
C
Candida Crawford
Third Space Learning, Swindon, UK.
T
Tom Hooper
Third Space Learning, Swindon, UK.
Gautam Biswas
Gautam Biswas
Cornelius Vanderbilt Professor of Engineering; Professor of Computer Science and Engineering
Artificial Intelligence in EducationReinforcement LearningMultimodal AnalyticsLearning Tech
R
Rene Kizilcec
College of Computing and Information Science, Cornell University, Ithaca, USA.