🤖 AI Summary
This study addresses the limitation of existing models that focus solely on image interpretation while lacking prognostic reasoning integrated with clinical history, thereby hindering opportunistic major adverse cardiovascular event (MACE) risk prediction. To overcome this, we propose a causal reinforcement learning framework that integrates chest radiographs and medical histories for multimodal clinical reasoning. Specifically, we introduce a novel role-decoupled dual-LLM architecture that separates reasoning from prediction, alongside a dual-action causal reinforcement strategy and causal token pruning to optimize evidence selection and representation compression. This work achieves a paradigm shift from image interpretation to clinical reasoning. Evaluated across multiple datasets, the proposed method attains AUROCs of 0.72–0.845, significantly outperforming baselines while substantially improving both reasoning quality and expert preference.
📝 Abstract
Major adverse cardiovascular events (MACE) remain the leading cause of mortality worldwide. Opportunistic screening using routinely acquired clinical data offers a scalable approach for identifying high-risk individuals before acute events occur. Although chest X-rays (CXRs) capture latent cardiovascular biomarkers and clinical histories provide complementary patient context, existing medical vision-language models are primarily optimized for radiology interpretation rather than prognostic reasoning. We propose a causal reinforcement learning framework for multimodal clinical reasoning that integrates CXRs and physician-authored clinical histories for opportunistic MACE prediction. The framework introduces (1) a role-decoupled dual-LLM architecture that separates reasoning from risk prediction, (2) a dual-action causal reinforcement learning policy for evidence selection and reasoning optimization, and (3) causal token pruning to learn compact multimodal representations. Evaluated on an internal cohort, an emergency department cohort, and the external MIMIC dataset, the proposed framework consistently outperformed unimodal baselines and state-of-the-art medical vision-language models, achieving AUROCs of 0.720, 0.760, and 0.845, respectively. It also substantially improved reasoning quality, achieving higher GREEN scores and higher expert preference while maintaining robust predictive performance across diverse patient populations.