🤖 AI Summary
This study investigates whether score discrepancies in language model-based depression assessments stem from patient symptomatology or inherent model bias. Employing preregistered experiments and multi-model cross-validation, we conduct a large-scale quantitative analysis utilizing the PHQ-8 scale, ROC curves, and recalibration techniques. We provide the first empirical evidence that model selection explains substantially more scoring variance than individual differences, revealing significant "rater bias" in AI diagnostics. Notably, 40% of screening decisions are inconsistent across models, challenging conventional assumptions regarding AI assessment reliability. Although recalibration improves accuracy to 75%, substantial inter-model disagreement at the individual level persists, underscoring critical risks associated with clinical deployment.
📝 Abstract
Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals? We pre-registered 880 language-model raters, crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire. Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%. Two randomly drawn raters with area under the receiver operating characteristic curve (AUC) >= 0.70 disagreed on screening decisions for 40% of participants, on average. Average over-rating governed how many were flagged, yet equal-capacity raters chose differently for about one participant in five. A locked analysis of 86 new interviews reproduced the main pre-registered findings. Exploratory recalibration with 40 labelled participants raised accuracy from about 60% to 75% and halved disagreement, leaving one participant in five decided differently. Calibration repaired much of the rater dependence without securing agreement about individuals.