π€ AI Summary
This study addresses the verification bias in medical report auditing caused by omissions in AI-generated checklists. To mitigate this issue, we propose training a vision-language model via reinforcement learning to populate comprehensive checklists, which are subsequently used by an independent model to perform automated report verification. Our investigation reveals that data formatting and label consistency significantly influence the modelβs discriminative capacity and false positive rate. Experimental results demonstrate that the proposed approach improves the Youden index by over 11%. Furthermore, the analysis highlights substantial discrepancies among different verifiers in their acceptance rates of negative statements. These findings provide critical insights for optimizing automated medical report verification systems.
π Abstract
Automated checks of radiology reports may rely on AI-generated checklists that leave findings unmentioned. We used reinforcement learning to train a vision-language model to fill in a 12-finding checklist from a chest radiograph without seeing the sentence under test; a separate checking model judged the sentence from the checklist. On held-out patients, a rule-based check and an independent medical checker, neither used in training, measured discrimination gains (Youden index) of 12.6% and 11.8%; only the rule-based check met the prespecified false-alarm criterion. Switching to the training format, which fixes finding order and enters unmentioned findings as absent, raised the training checker's measured gain and lowered the independent checker's, a prespecified comparison that yielded 6.2% (95% interval 2.0% to 10.5%) and, post hoc on held-out patients, 7.7%. Across 8 checking models, acceptance of a label-consistent negative statement about an unmentioned finding ranged from 1.0% to 97.0%. Labels were report-derived, not radiologist-adjudicated.