Separating Capability from Confidence: Grounded Dual-State Calibration for GRPO-Trained Medical Vision-Language Models

📅 2026-09-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出DualRead方法,通过分离能力学习与置信度估计,提高医学视觉-语言模型的正确性判别和校准,同时保持答案准确性。
📝 Abstract
Medical vision-language models (VLMs) require confidence that reflects both answer correctness and patient-specific visual evidence. Recent GRPO-based methods optimize verbalized confidence together with answer generation. However, this joint optimization may interfere with answer learning and drive confidence toward near-binary values. Verbalized confidence also provides no explicit assessment of visual support. We therefore separate capability learning from confidence estimation and propose \textbf{DualRead}. DualRead builds on the insight that reliability can be read from the actor's internal states at critical moments in the answering process. It freezes the GRPO-trained actor and combines pre-answer solvability with a post-answer assessment of the generated answer and its visual support. To further assess whether confidence reflects visual grounding, we introduce \textbf{Counterfactual Confidence Grounding AUC} (CCG-AUC). It measures whether confidence decreases when real-image substitution changes the actor from correct to incorrect. Across two VLM backbones and both in- and out-of-distribution medical VQA benchmarks, DualRead improves correctness discrimination and calibration over verbalized confidence while preserving answer accuracy. CCG-AUC reveals whether confidence responds to answer-relevant visual evidence rather than primarily to non-visual cues.
Problem

Research questions and friction points this paper is trying to address.

Medical Vision-Language Models
Confidence Estimation
Visual Evidence
GRPO-Trained Models
Calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

DualRead
GRPO
Confidence Estimation
Visual Support
CCG-AUC
Y
Yangyang Xie
Shanghai Jiao Tong University, Shanghai, China
K
Ke Hao
Shanghai Jiao Tong University, Shanghai, China
J
Jiaqi Liu
Medical Image Insights Co. Ltd., Shanghai, China
Yun Gu
Yun Gu
Shanghai Jiao Tong University
Medical Image AnalysisComputer-Assisted Intervention
X
Xinglin Zhang
Medical Image Insights Co. Ltd., Shanghai, China