Rubric Spans are Label Representations: Joint LLM Encoding for Short Answer Scoring

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Automated short-answer grading faces persistent challenges in efficiency, cross-rubric transferability, and precise alignment with scoring criteria. This study proposes RUSPAN, a framework that models grading rubrics as semantic label representations, enabling listwise parallel scoring through single-pass joint encoding with large language models. Furthermore, it introduces the RUSPAN-RIM mechanism, which employs independent attention masks to mitigate overfitting, thereby substantially enhancing zero-shot cross-lingual and cross-structural transfer capabilities. Experiments conducted across six multilingual benchmarks demonstrate that the proposed approach consistently outperforms both discriminative and generative baselines, exhibiting particularly pronounced advantages in challenging transfer scenarios.
📝 Abstract
Automatic Short Answer Scoring (ASAS) requires models that can score student responses against question-specific criteria while remaining efficient and transferable across rubric sets. We propose RUSPAN, a rubric-conditioned ASAS framework that treats rubric descriptions as semantic label representations. RUSPAN serialises the question context, student answer, and all candidate rubric levels into a single sequence, then scores the levels listwise from the rubric-span and whole-sequence representations produced in a single LM pass. We further introduce RUSPAN-RIM, in which a Rubric-Independent Mask prevents rubric spans from attending to one another, making rubric representations depend only on the answer and question context and preventing overfitting to rubric patterns during training for zero-shot transfer. On six ASAS benchmarks spanning English, German, and Portuguese, RUSPAN improves mono-benchmark scoring over discriminative and generative baselines, while RIM with position reindexing delivers consistent and substantial gains on PT-ASAG, the held-out benchmark with the strongest combined language and rubric-structure shift.
Problem

Research questions and friction points this paper is trying to address.

Automatic Short Answer Scoring
Rubric Representation
Zero-shot Transfer
Cross-lingual Transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automatic Short Answer Scoring
Rubric Spans
Listwise Scoring
Rubric-Independent Mask
Zero-shot Transfer
💼 Related Jobs
No related jobs found.
Z
Zhifan Sun
DIPF | Leibniz Institute for Research and Information in Education
S
Sebastian Gombert
DIPF | Leibniz Institute for Research and Information in Education
Fabian Zehner
Fabian Zehner
DIPF | Leibniz Institute for Research and Information in Education, Centre for International Student Assessment (ZIB)
L
Leon Camus
DIPF | Leibniz Institute for Research and Information in Education
L
Longwei Cong
DIPF | Leibniz Institute for Research and Information in Education
Hendrik Drachsler
Hendrik Drachsler
Professor for Computer Science, DIPF | Leibniz Institute & Goethe University, Frankfurt
Learning AnalyticsAI in EducationAssessment and FeedbackLearning DesignMedical Education