🤖 AI Summary
This study addresses the failure of existing datasets caused by mixed text-image-formula layouts and erasure noise in real-world student answer sheets by constructing HANS, the first hybrid document parsing benchmark tailored for educational scenarios. Furthermore, this work proposes NA-GOT, an end-to-end framework that achieves robust recognition under complex noise through fine-grained annotation, a lightweight feature denoising module, and a noise-aware attention mechanism during decoding. Experimental results demonstrate that the proposed dataset presents significant challenges, while NA-GOT yields substantial improvements in both the accuracy and stability of answer process recognition.
📝 Abstract
Intelligent grading and automated scoring technologies constitute critical infrastructure for smart education. However, existing document parsing and handwriting recognition benchmarks are predominantly designed for well-structured printed documents or isolated mathematical expressions, lacking datasets that capture the complex characteristics inherent to student answer sheets, including multi-line derivation processes, heterogeneous mixtures of text and mathematical formulae, and noise artifacts such as strikethroughs. To address this gap, we introduce HANS, the first dataset explicitly constructed for real-world educational scenarios, encompassing mathematical expressions, natural language text, hand-drawn tables, and diverse noise patterns including corrections and deletions, accompanied by fine-grained annotations that establish a reliable foundation for robust recognition research. Building upon HANS, we propose NA-GOT, an end-to-end framework that achieves two-stage noise suppression through a lightweight noise suppression module operating at the feature level, complemented by a noiseaware attention mechanism incorporated into the decoding stage. Experimental results demonstrate that HANS poses substantial challenges to existing methods, while NA-GOT achieves significant improvements in both accuracy and stability for answer process recognition. The dataset will be made publicly available upon publication.