An Information-Theoretic Evaluation Framework for Benchmark and Model Diagnosis in Knowledge Tracing
This study addresses the limitations of global evaluation metrics, which obscure error sources and complicate the assessment of benchmark saturation, by proposing an information-theoretic fine-grained evaluation framework. It introduces the Context Tree Weighting (CTW) algorithm as an operational coordinate to quantify local predictability at the entropy-band level, effectively distinguishing causal from irreducible uncertainty. Experiments demonstrate that modern models achieve significant yet non-uniform gains in high-entropy regions. Furthermore, the proposed method precisely identifies residual predictive structures and noise-sensitive areas, revealing inherent benchmark limitations. Ultimately, this framework provides a reliable analytical tool for diagnosing the performance of large language models.