🤖 AI Summary
This study investigates whether large language models (LLMs) can accurately simulate the cognitive biases of novice programmers. Through code-tracing tasks integrated with dynamic execution analysis, it establishes a multidimensional benchmark evaluating thirteen LLMs in generating, solving, and diagnosing misconceptions. The findings reveal that while frontier models perform reliably, smaller models struggle with dynamic execution. Furthermore, fine-tuned models tend to revert to correct reasoning, exposing a "curse of knowledge" effect. Additionally, true-or-false question formats facilitate more precise diagnosis than open-ended generation. Overall, this work highlights significant limitations of current LLMs in authentically simulating novice errors, providing critical insights for the development of educational AI systems.
📝 Abstract
We evaluate 13 LLMs on generating, solving, simulating, and diagnosing novice programming misconceptions on code-tracing problems. While frontier models reliably perform these tasks, small models ($\le$14B) struggle with dynamic execution. Notably, code-tuned models fail during misconception simulation by reverting to correct execution. Finally, misconception diagnosis is significantly easier in True/False formats than in open-ended generation tasks.