From Image Interpretation to Clinical Reasoning: Upstream Physician-Context-Aware Multimodal Learning with Causal Reinforcement Learning

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing models that focus solely on image interpretation while lacking prognostic reasoning integrated with clinical history, thereby hindering opportunistic major adverse cardiovascular event (MACE) risk prediction. To overcome this, we propose a causal reinforcement learning framework that integrates chest radiographs and medical histories for multimodal clinical reasoning. Specifically, we introduce a novel role-decoupled dual-LLM architecture that separates reasoning from prediction, alongside a dual-action causal reinforcement strategy and causal token pruning to optimize evidence selection and representation compression. This work achieves a paradigm shift from image interpretation to clinical reasoning. Evaluated across multiple datasets, the proposed method attains AUROCs of 0.72–0.845, significantly outperforming baselines while substantially improving both reasoning quality and expert preference.
📝 Abstract
Major adverse cardiovascular events (MACE) remain the leading cause of mortality worldwide. Opportunistic screening using routinely acquired clinical data offers a scalable approach for identifying high-risk individuals before acute events occur. Although chest X-rays (CXRs) capture latent cardiovascular biomarkers and clinical histories provide complementary patient context, existing medical vision-language models are primarily optimized for radiology interpretation rather than prognostic reasoning. We propose a causal reinforcement learning framework for multimodal clinical reasoning that integrates CXRs and physician-authored clinical histories for opportunistic MACE prediction. The framework introduces (1) a role-decoupled dual-LLM architecture that separates reasoning from risk prediction, (2) a dual-action causal reinforcement learning policy for evidence selection and reasoning optimization, and (3) causal token pruning to learn compact multimodal representations. Evaluated on an internal cohort, an emergency department cohort, and the external MIMIC dataset, the proposed framework consistently outperformed unimodal baselines and state-of-the-art medical vision-language models, achieving AUROCs of 0.720, 0.760, and 0.845, respectively. It also substantially improved reasoning quality, achieving higher GREEN scores and higher expert preference while maintaining robust predictive performance across diverse patient populations.
Problem

Research questions and friction points this paper is trying to address.

Major adverse cardiovascular events
Multimodal clinical reasoning
Medical vision-language models
Opportunistic screening
Prognostic reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Reinforcement Learning
Multimodal Clinical Reasoning
Role-Decoupled Dual-LLM
Causal Token Pruning
MACE Prediction
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
J
Jialu Pi
Dept. of Radiology, Mayo Clinic, Phoenix, AZ, USA; School of Computing and Augmented Intelligence, Arizona State University, Tempe, USA
Yanan Ma
Yanan Ma
City University of Hong Kong
Wireless networksEdge intelligence
W
Weijie Chen
Dept. of Radiology, Mayo Clinic, Phoenix, USA
O
Owen Crystal
Dept. of Radiology, Mayo Clinic, Phoenix, USA
S
Shubham Trivedi
Dept. of Radiology, Mayo Clinic, Phoenix, USA
S
Stephen Xie
Mayo Clinic, Phoenix, USA
A
Anna Silverman
Mayo Clinic, Phoenix, USA
M
Matthew Stib
Dept. of Radiology, Mayo Clinic, Phoenix, USA
C
Chadi Ayoub
Dept. of Cardiology, Mayo Clinic, Phoenix, USA
Reza Arsanjani
Reza Arsanjani
Mayo Clinic Arizona
Echocardiac imagingvalvescardio-oncology
Imon Banerjee
Imon Banerjee
Mayo Clinic, AZ
Deep learningNatural language processingPredictive modeling3D characterizationMedical image annotation