🤖 AI Summary
Automated short-answer grading faces persistent challenges in efficiency, cross-rubric transferability, and precise alignment with scoring criteria. This study proposes RUSPAN, a framework that models grading rubrics as semantic label representations, enabling listwise parallel scoring through single-pass joint encoding with large language models. Furthermore, it introduces the RUSPAN-RIM mechanism, which employs independent attention masks to mitigate overfitting, thereby substantially enhancing zero-shot cross-lingual and cross-structural transfer capabilities. Experiments conducted across six multilingual benchmarks demonstrate that the proposed approach consistently outperforms both discriminative and generative baselines, exhibiting particularly pronounced advantages in challenging transfer scenarios.
📝 Abstract
Automatic Short Answer Scoring (ASAS) requires models that can score student responses against question-specific criteria while remaining efficient and transferable across rubric sets. We propose RUSPAN, a rubric-conditioned ASAS framework that treats rubric descriptions as semantic label representations. RUSPAN serialises the question context, student answer, and all candidate rubric levels into a single sequence, then scores the levels listwise from the rubric-span and whole-sequence representations produced in a single LM pass. We further introduce RUSPAN-RIM, in which a Rubric-Independent Mask prevents rubric spans from attending to one another, making rubric representations depend only on the answer and question context and preventing overfitting to rubric patterns during training for zero-shot transfer. On six ASAS benchmarks spanning English, German, and Portuguese, RUSPAN improves mono-benchmark scoring over discriminative and generative baselines, while RIM with position reindexing delivers consistent and substantial gains on PT-ASAG, the held-out benchmark with the strongest combined language and rubric-structure shift.