MedQA-MM: Shortcuts Behind Medical Visual Reasoning

📅 2026-09-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文通过多种方法分析并解决医疗多模态选择题中因非图像线索导致的推理膨胀问题,构建了减少捷径依赖的数据集MedQA-MM。
📝 Abstract
A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the intended image finding or by benchmark-preserved cues in the wording of answers, non-visual clinical text, visible image text, artificial annotations, or device/context artifacts. We call the resulting score-level overinterpretation reasoning inflation. Here, a route is an observable input path that can support answer selection, not a claim about the model's hidden cognition. Across six medical multimodal MCQ datasets, we separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key. In a 13-configuration open-model panel, full-input accuracy is 62.63%, while text-only and options-only settings achieve 53.96% and 29.71%, respectively. Removing length-gap, absolute/conspicuous, and spatial/prepositional cues lowers accuracy by 6.58, 3.50, and 4.77 percentage points. We also construct MedQA-MM, a 1,000-item shortcut-mitigated subset, where text-only and options-only accuracy fall to 5.21% and 12.33%. This does not imply that models never use images; it shows that medical image-reasoning claims require route-level evidence.
Problem

Research questions and friction points this paper is trying to address.

Medical Multimodal MCQs
Reasoning Inflation
Benchmark Score
Visual Reasoning
Shortcut Cues
Innovation

Methods, ideas, or system contributions that make the work stand out.

Medical Visual Reasoning
Benchmark Score
Reasoning Inflation
Route-Level Evidence
Shortcut Mitigation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Benlu Wang
University of Massachusetts Amherst
Y
Yifan Zhang
University of Massachusetts Lowell
J
Jiaqing Yu
Qingdao Medical College of Qingdao University
C
Chin Siang Ong
Yale School of Medicine
J
Juncheng Huang
National University Hospital, Singapore
Z
Zhuohao Li
Zhejiang University
Z
Zhenyu Zhang
Stanford University
Arman Cohan
Arman Cohan
Yale University; Allen Institute for AI
Natural Language ProcessingMachine LearningArtificial Intelligence
H
Hong Yu
University of Massachusetts Amherst
Zonghai Yao
Zonghai Yao
Umass Amherst
Medical-LLMMulti-agent AI HospitalClinical ReasoningSynthetic DataPatient Education