Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of large language model (LLM)-based automated scoring to prompt sensitivity, which frequently induces bias or failure and impedes its reliable substitution for scarce human evaluation. By systematically assessing LLM scoring performance across 171 configurations, this work reveals a fragility mechanism wherein overly stringent prompts precipitate model collapse, and proposes a lightweight LoRA fine-tuning strategy to restore robustness and accuracy. Validated through double-blind experiments and ablation analyses, the proposed approach enables five open-source models to achieve expert-level agreement after fine-tuning, with the optimal model attaining a mean absolute error (MAE) superior to human inter-rater consistency. The complete dataset and code have been publicly released to facilitate reproducibility.
📝 Abstract
One long-form exam in a large course costs hundreds of grader-hours, and qualified graders are scarce; LLM graders are a tempting alternative. To show its pitfalls we grade a practical Computer Vision exam ($570$ dual-graded students) under $171$ configurations spanning closed and open-weights models; the best reaches mean absolute error $1.64/35$, below the $2.61/35$ two human graders achieve against each other. The catch is the prompt: a short ''strict grader'' preamble drives $14$ of $17$ open-weights models out of the graded band ($\text{MAE} \ge 8$), three stopping grading altogether. The damage traces to the preamble's two credit-withholding sentences, not to tone or model scale; one of them, ''never give partial credit'', alone makes two of three probed models stop grading. The closed flagships of three vendors shift calibration under it but stay in the band. In $162$ further configurations on a second, independent Machine Learning exam from another course ($1{,}038$ dual-graded students), the preamble worsens ten models, moving three out of the band into collapse and one into refusal, yet improves seven whose neutral prompts over-mark: the vulnerability replicates, but its direction is exam-specific. Light LoRA fine-tuning repairs it: one adapter on the two exams' pooled $\sim 3{,}900$ graded examples brings five small open models to parity or better with a human grader in agreement with the grader pair, and sensitivity to the three harsh personas nearly vanishes ($\le 0.32$ MAE). We release the anonymised dataset, full ablation grid, and grading, fine-tuning and analysis pipelines.
Problem

Research questions and friction points this paper is trying to address.

LLM graders
automated grading
prompt sensitivity
model failure
computer science exams
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM grading
prompt sensitivity
open-weights models
LoRA fine-tuning
grader calibration
A
Ali Habibullah
KAUST Academy, Computer, Electrical & Mathematical Sciences & Engineering (CEMSE), King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
Y
Yazan Alshoibi
KAUST Academy, Computer, Electrical & Mathematical Sciences & Engineering (CEMSE), King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
M
Mohammad Alshiekh
KAUST Academy, Computer, Electrical & Mathematical Sciences & Engineering (CEMSE), King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
Salman Khan
Salman Khan
Research Fellow, Oxford Brookes University
Computer VisionDeep learningFire / Smoke detectionAction RecognitionMedical Image Analysis
N
Naeemullah Khan
KAUST Academy & CEMSE, KAUST; Lady Margaret Hall, University of Oxford