An Integrated Machine Learning and Hierarchical Variance Decomposition Pipeline for Student Performance Prediction and Metacognitive Calibration on Multi-Signal Telemetry

📅 2026-06-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the longstanding disconnect between student academic performance prediction and metacognitive calibration by proposing a Unified Behavioral Prediction and Calibration Analysis Pipeline (UBP-CAP). Integrating prediction, calibration assessment, and variance decomposition modules, UBP-CAP leverages multimodal telemetry data to simultaneously predict response accuracy and quantify metacognitive bias. The work introduces the Prediction-Explanation Discrepancy Index (PEDI) to measure feature consistency between predictive and explanatory models and employs cross-validated generalized linear mixed-effects models (GLMMs) to uncover the context-dependence of calibration bias. Empirical results show that logistic regression (AUC = 0.903) outperforms LightGBM; students exhibit significantly higher calibration error (ECE = 0.109) than the model (ECE = 0.068); GLMM analysis yields an intraclass correlation coefficient (ICC) of 0.123, indicating calibration is predominantly context-driven; and PEDIcos = 0.081 reveals high alignment between prediction and explanation features.
📝 Abstract
Predicting student performance and characterizing metacognitive calibration are essential for personalization in intelligent tutoring systems. Prior research treats performance prediction, calibration error calculation, and variance decomposition as separate pipelines, preventing unified interpretation. I propose the Unified Behavioral Prediction and Calibration Analysis Pipeline (UBP-CAP), an integrated framework processing student pre-execution behavioral telemetry through three linked modules: (1) a LightGBM classifier with SHAP for binary correctness prediction, (2) formal calibration metrics (ECE, MCE, and Brier score decomposition) to evaluate metacognitive alignment, and (3) a crossed Generalized Linear Mixed-Effects Model (GLMM) for decomposing calibration deviations. I introduce the Predictive-Explanatory Divergence Index (PEDI), which quantifies structural divergence between predictive and explanatory feature profiles. Evaluated on 1,195 interaction records (27 students, 45 tasks), Logistic Regression achieves AUC-ROC = 0.903, outperforming LightGBM (0.878). Student naive ECE (0.109) significantly exceeds model ECE (0.068), confirming systematic miscalibration. The crossed GLMM yields ICCStudent = 0.123, showing calibration is situational rather than dispositional. PEDIcos = 0.081 (p = 0.327) indicates structural alignment between prediction and explanation on shared behavioral features.
Problem

Research questions and friction points this paper is trying to address.

student performance prediction
metacognitive calibration
variance decomposition
intelligent tutoring systems
calibration error
Innovation

Methods, ideas, or system contributions that make the work stand out.

metacognitive calibration
variance decomposition
PEDI
GLMM
SHAP
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Gurdeep Singh Virdee
Fergana State Technical University, Uzbekistan