AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their Impact on Automated Assessment

📅 2026-08-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical limitation in existing rubric-based AI systems for evaluating academic papers: their reliance on predefined weights lacking empirical grounding in instructors’ actual assessment preferences, which contributes to discrepancies between AI-generated and human scores. To bridge this gap, the authors surveyed 84 interdisciplinary instructors to elicit their assigned weights across 35 evaluation criteria, integrating these empirically derived weights into the AI system RubiSCoT. The impact of different weighting schemes on scoring alignment was systematically evaluated using 80 German-language theses. Despite being the first large-scale empirical calibration of AI evaluation weights, the optimal configuration reduced the mean relative deviation between AI and instructor scores only marginally—from 11.18% to 10.85%—while inter-instructor deviation remained substantially lower at 4.44%, suggesting that weight adjustment alone is insufficient to achieve meaningful alignment between AI and human grading.
📝 Abstract
Rubric-based AI systems for thesis assessment use criterion weights to assign different levels of importance to evaluation criteria. These weights are typically defined through expert judgment, although little empirical evidence exists regarding how thesis supervisors actually prioritize evaluation criteria. Consequently, this study investigates supervisor-derived criterion weights in thesis assessment and evaluates their impact on AI-based assessment. We surveyed 84 thesis supervisors across four academic disciplines and collected weighting data for 35 thesis assessment criteria. Comparison with the default criterion weights of the AI assessment system RubiSCoT [1] revealed substantial divergences between supervisor-derived and default criterion weights. To evaluate the practical implications of these differences, the supervisor-derived weights were integrated into multiple calibration configurations and evaluated on a corpus of 80 German-language theses. The best-performing configuration reduced the mean relative deviation between AI-generated and supervisor-assigned evaluations from 11.18% to 10.85%, although the improvement was not statistically significant. Human supervisors showed substantially stronger agreement with each other, exhibiting a mean inter-supervisor relative deviation of 4.44%. The findings indicate that criterion-weight calibration alone does not substantially improve alignment between AI-generated and human assessments.
Problem

Research questions and friction points this paper is trying to address.

AI-based assessment
thesis evaluation
criterion weights
human-AI alignment
rubric calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

criterion-weight calibration
empirical study
AI-based thesis assessment
human-AI alignment
rubric-based evaluation
🔎 Similar Papers