Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of public benchmarks and the difficulty of fine-grained evaluation in automated short-answer grading for German. To this end, we construct a large-scale, rubric-based German dataset encompassing three subtasks: learning performance, knowledge components, and skills. Methodologically, we propose a novel paradigm that formulates grading as a retrieval task, enabling fine-grained assessment aligned with multidimensional educational competencies. We further conduct extensive multi-model benchmarking using zero-shot prompting with large language models (LLMs), encoders, and classifiers. Our findings reveal that LLMs exhibit limitations in evaluating knowledge and skills under zero-shot settings, whereas incorporating explicit rubric texts significantly enhances assessment performance.
📝 Abstract
Automatic Short Answer Scoring (ASAS) is central to NLP for Education. However, openly available benchmarks remain scarce, and existing datasets largely address how well students answer a question directly rather than how well they master underlying concepts (knowledge elements) such as thermal energy or epistemic activities (skills) such as reasoning or claim. To address this gap, we introduce Alice, a large-scale, rubric-based German ASAS dataset that is pedagogically aligned and comprises three subtasks: (i) learning performance (Alice-LP), (ii) knowledge elements (Alice-KE), and (iii) skills (Alice-SK). We further formulate rubric-based ASAS as a rubric-retrieval task and benchmark the dataset with a range of language models, from encoder-only models to lightweight LLMs. We also benchmark the dataset with zero-shot prompting via LLMs and a standard classification baseline. The experiments show that LLMs, in particular, struggle to score knowledge elements and skills in the zero-shot setting. They also indicate that rubric text is often useful, especially for Alice-KE and Alice-SK, while on Alice-LP gains over sample-solution-focused inputs are more modest and vary by model and input format.
Problem

Research questions and friction points this paper is trying to address.

Automatic Short Answer Scoring
Benchmark Dataset
Rubric-Based Assessment
Knowledge Elements
Skills
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automatic Short Answer Scoring
Rubric-based Evaluation
Rubric Retrieval
Multi-dimensional Scoring
Benchmark Dataset
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhifan Sun
DIPF | Leibniz Institute for Research and Information in Education
S
Sebastian Gombert
DIPF | Leibniz Institute for Research and Information in Education
J
Jannik Lossjew
IPN | Leibniz Institute for Science and Mathematics Education
T
Tobias Wyrwich
IPN | Leibniz Institute for Science and Mathematics Education
B
Berrit Katharina Czinczel
IPN | Leibniz Institute for Science and Mathematics Education
D
David Bednorz
IPN | Leibniz Institute for Science and Mathematics Education
M
Marcus Kubsch
Umeå University
Knut Neumann
Knut Neumann
Leibniz Institut für die Pädagogik der Naturwissenschaften und der Mathematik
Hendrik Drachsler
Hendrik Drachsler
Professor for Computer Science, DIPF | Leibniz Institute & Goethe University, Frankfurt
Learning AnalyticsAI in EducationAssessment and FeedbackLearning DesignMedical Education