Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring
This study addresses the lack of public benchmarks and the difficulty of fine-grained evaluation in automated short-answer grading for German. To this end, we construct a large-scale, rubric-based German dataset encompassing three subtasks: learning performance, knowledge components, and skills. Methodologically, we propose a novel paradigm that formulates grading as a retrieval task, enabling fine-grained assessment aligned with multidimensional educational competencies. We further conduct extensive multi-model benchmarking using zero-shot prompting with large language models (LLMs), encoders, and classifiers. Our findings reveal that LLMs exhibit limitations in evaluating knowledge and skills under zero-shot settings, whereas incorporating explicit rubric texts significantly enhances assessment performance.