Seeing the Big Picture: Evaluating Multimodal LLMs' Ability to Interpret and Grade Handwritten Student Work

📅 2025-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates multimodal large language models (MLLMs) for automated parsing and grading of elementary school handwritten mathematics assignments, focusing on arithmetic answer recognition and open-ended mathematical diagram evaluation. To mitigate MLLMs’ overreliance on image quality, we propose a novel “description-augmented” paradigm that incorporates human-generated semantic descriptions as auxiliary textual input alongside visual inputs. Experimental results show that, on objective arithmetic problems, MLLMs achieve 95% accuracy (Cohen’s κ = 0.90) when processing raw images—approaching human-level performance. In contrast, for diagram evaluation, inter-rater agreement drops to κ = 0.20 with image-only input but improves significantly to κ = 0.47 with description augmentation—matching the consistency observed among human graders. This work is the first to empirically reveal a substantial capability gap between MLLMs’ performance on objective versus open-ended handwritten mathematics tasks, and demonstrates that text-based augmentation substantially enhances both reliability and interpretability in open-ended educational assessment.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsComputer Vision: Large Vision Models

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
Recent advances in multimodal large language models (MLLMs) raise the question of their potential for grading, analyzing, and offering feedback on handwritten student classwork. This capability would be particularly beneficial in elementary and middle-school mathematics education, where most work remains handwritten, because seeing students' full working of a problem provides valuable insights into their learning processes, but is extremely time-consuming to grade. We present two experiments investigating MLLM performance on handwritten student mathematics classwork. Experiment A examines 288 handwritten responses from Ghanaian middle school students solving arithmetic problems with objective answers. In this context, models achieved near-human accuracy (95%, k = 0.90) but exhibited occasional errors that human educators would be unlikely to make. Experiment B evaluates 150 mathematical illustrations from American elementary students, where the drawings are the answer to the question. These tasks lack single objective answers and require sophisticated visual interpretation as well as pedagogical judgment in order to analyze and evaluate them. We attempted to separate MLLMs' visual capabilities from their pedagogical abilities by first asking them to grade the student illustrations directly, and then by augmenting the image with a detailed human description of the illustration. We found that when the models had to analyze the student illustrations directly, they struggled, achieving only k = 0.20 with ground truth scores, but when given human descriptions, their agreement levels improved dramatically to k = 0.47, which was in line with human-to-human agreement levels. This gap suggests MLLMs can "see" and interpret arithmetic work relatively well, but still struggle to "see" student mathematical illustrations.
Problem

Research questions and friction points this paper is trying to address.

Evaluating MLLMs' ability to grade handwritten student work
Assessing MLLM performance on mathematical illustrations and arithmetic
Investigating gaps in visual interpretation versus pedagogical judgment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluating MLLMs for grading handwritten student work
Testing models on arithmetic answers and mathematical illustrations
Augmenting images with human descriptions to improve accuracy
🔎 Similar Papers
No similar papers found.
O
Owen Henkel
University of Oxford
B
Bill Roberts
Legible Labs
D
Doug Jaffe
Coherence Fund
L
Laurence Holt
XQ Institute