Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究使用Jev模型作为低成本判断工具,评估AI生成的放射学报告与医生报告之间的一致性及事实差异。
📝 Abstract
An AI-generated radiology report can resemble a physician's report while omitting an abnormality, adding an unsupported finding, or reversing its presence. Measuring these factual differences is essential for evaluating report generators. We study Jev, a System One decision model, as a simple, low-cost judge of agreement with physician-written reference reports. Our evaluator checks whether each statement is supported by the other report and combines these judgments in both directions to capture unsupported claims and omissions. A single-question configuration reaches Kendall correlations of 0.573 on RadEvalX and 0.398 on RadEvalExpert with expert error counts, outperforming an open natural language inference judge under matched decomposition and aggregation. One support question per statement retains similar expert agreement to seven while using 43-45% fewer judgment input tokens. At the documented API price, judgments cost under three cents per hundred report pairs, excluding local decomposition. In a separate controlled-error test, Jev detects false negation with an AUROC of 0.977. Local RadMatch achieves stronger agreement on clinically significant errors in both expert datasets and on total errors in the shared RadEvalExpert subset. Finding-count and error-scope analyses show that benchmark agreement reflects report size and error definitions as well as medical error detection. These results support Jev as a practical judgment component for measuring factual differences in generated radiology reports and identify where more elaborate evaluation remains valuable.
Problem

Research questions and friction points this paper is trying to address.

Radiology Reports
Clinical Factuality
AI-generated
Factual Differences
Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Jev
System One decision model
Radiology reports evaluation
Factual differences measurement
Cost-effective
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jiaju Huang
Intelligent Medical Computing Laboratory, Faculty of Applied Sciences, Macao Polytechnic University
H
Hao Yang
Intelligent Medical Computing Laboratory, Faculty of Applied Sciences, Macao Polytechnic University
X
Xinyu Ma
Intelligent Medical Computing Laboratory, Faculty of Applied Sciences, Macao Polytechnic University
X
Xinglong Liang
Department of Radiology, Netherlands Cancer Institute, Amsterdam, the Netherlands; Department of Radiology and Nuclear Medicine, Radboud University Medical Center, Nijmegen, the Netherlands
K
Kunyan Cai
Intelligent Medical Computing Laboratory, Faculty of Applied Sciences, Macao Polytechnic University
J
Junqiang Ma
Intelligent Medical Computing Laboratory, Faculty of Applied Sciences, Macao Polytechnic University
S
Shaobin Chen
Intelligent Medical Computing Laboratory, Faculty of Applied Sciences, Macao Polytechnic University
Yue Sun
Yue Sun
Macao Polytechnic University; Eindhoven University of Technology
Video ProcessingImage ProcessingMachine LearningAI in Healthcare
Tao Tan
Tao Tan
FCA MPU
Medical Imaging AI