Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study
This study addresses the risk of erroneous AI recommendations misleading junior physicians in AI-assisted diagnosis by conducting a prospective, multicenter randomized controlled trial comparing dual-model versus single-model AI support for radiographic interpretation. Utilizing the multimodal large language models GPT-5.4, Kimi-K2.6, and Gemini-3.6 Flash, this work provides the first quantitative evaluation of the error-correction capability inherent in a dual-model cross-validation mechanism, with statistical analyses performed using Welch’s ANOVA and HC3 robust linear regression. Results demonstrate that dual-recommendation support significantly mitigates the misleading influence of incorrect AI outputs, improving diagnostic accuracy among radiologists by approximately seven percentage points. Conversely, non-radiologists derived no significant benefit, revealing a pronounced interaction effect driven by specialty background.