🤖 AI Summary
This study addresses the lack of anatomical localization capabilities and clinical question-answering supervision data in vision-language models for cardiac MRI by constructing a large-scale annotated dataset and proposing the CARA-VL model. Its core innovation lies in introducing the first Context-Aware Anatomical Routing Attention (CARA) mechanism, which achieves precise anatomical guidance by dynamically selecting anatomical priors and modulating decoder attention intensity. Methodologically, the framework integrates anatomy-grounded pretraining with task-specific attention strategies. Experimental results demonstrate that the proposed model significantly improves regional localization accuracy and clinical evaluation performance while exhibiting strong generalization capability across external cohorts.
📝 Abstract
Cardiac magnetic resonance imaging (CMR) enables assessment of cardiac anatomy, ventricular function, and myocardial tissue characteristics. Clinicians interpret these images by identifying cardiac structures and focusing on the regions relevant to each clinical question, motivating anatomically guided vision-language models (VLMs). Yet CMR-specific supervision for anatomical localisation and clinical question answering remains limited. To address this gap, we investigate fine-grained CMR visual question answering through anatomical grounding and guided attention. We construct 128,915 anatomical-grounding and 42,799 clinical QA pairs across short-axis cine, late gadolinium enhancement, and long-axis cine. These datasets support anatomical recognition, localisation, and clinical assessment without requiring paired reports for individual training images. To help the model learn where to look, we introduce Cardiac Anatomy-Routed Attention (CARA), which selects predicted anatomical priors according to the question and guides decoder attention with learned task-specific strengths. Combining anatomical grounding pretraining with CARA yields our model, CARA-VL. Experiments demonstrate CARA-VL's strengths in clinical assessment and regional localisation across CMR imaging settings, with promising generalization to an external clinical cohort. Together, our data and method provide a practical framework for studying and advancing cardiac visual understanding in VLMs. We will release the QA data derived from public datasets upon publication.