Institution profile

Leibniz Institute for Science and Mathematics Education

Academic institutioneurope · de
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring

Oct 07, 2026

This study addresses the lack of public benchmarks and the difficulty of fine-grained evaluation in automated short-answer grading for German. To this end, we construct a large-scale, rubric-based German dataset encompassing three subtasks: learning performance, knowledge components, and skills. Methodologically, we propose a novel paradigm that formulates grading as a retrieval task, enabling fine-grained assessment aligned with multidimensional educational competencies. We further conduct extensive multi-model benchmarking using zero-shot prompting with large language models (LLMs), encoders, and classifiers. Our findings reveal that LLMs exhibit limitations in evaluating knowledge and skills under zero-shot settings, whereas incorporating explicit rubric texts significantly enhances assessment performance.

0 citationsRead paper

Evidence for Daily and Weekly Periodic Variability in GPT-4o Performance

Feb 06, 2026arXiv.org

This study investigates whether the performance of large language models remains constant over time by conducting a longitudinal replication experiment on GPT-4o over three months, with ten evaluations every three hours under identical physics tasks. Employing time-series data collection, controlled experimental design, and Fourier spectral analysis, the research reveals—for the first time—significant diurnal and weekly periodic fluctuations in model performance, with approximately 20% of output variance attributable to these rhythmic patterns. These findings challenge the foundational assumption of temporal invariance in model performance, demonstrating that the output quality of large language models exhibits systematic time dependence. The results carry important implications for the reproducibility of AI research and the prevailing paradigms used to evaluate model capabilities.

0 citationsRead paper
Recent publications

Latest Papers

Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring

Oct 07, 2026

This study addresses the lack of public benchmarks and the difficulty of fine-grained evaluation in automated short-answer grading for German. To this end, we construct a large-scale, rubric-based German dataset encompassing three subtasks: learning performance, knowledge components, and skills. Methodologically, we propose a novel paradigm that formulates grading as a retrieval task, enabling fine-grained assessment aligned with multidimensional educational competencies. We further conduct extensive multi-model benchmarking using zero-shot prompting with large language models (LLMs), encoders, and classifiers. Our findings reveal that LLMs exhibit limitations in evaluating knowledge and skills under zero-shot settings, whereas incorporating explicit rubric texts significantly enhances assessment performance.

0 citationsRead paper

Evidence for Daily and Weekly Periodic Variability in GPT-4o Performance

Feb 06, 2026arXiv.org

This study investigates whether the performance of large language models remains constant over time by conducting a longitudinal replication experiment on GPT-4o over three months, with ten evaluations every three hours under identical physics tasks. Employing time-series data collection, controlled experimental design, and Fourier spectral analysis, the research reveals—for the first time—significant diurnal and weekly periodic fluctuations in model performance, with approximately 20% of output variance attributable to these rhythmic patterns. These findings challenge the foundational assumption of temporal invariance in model performance, demonstrating that the output quality of large language models exhibits systematic time dependence. The results carry important implications for the reproducibility of AI research and the prevailing paradigms used to evaluate model capabilities.

0 citationsRead paper