Learning Where to Look: Anatomical Grounding and Guided Attention for Cardiac MRI Vision-Language Models

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of anatomical localization capabilities and clinical question-answering supervision data in vision-language models for cardiac MRI by constructing a large-scale annotated dataset and proposing the CARA-VL model. Its core innovation lies in introducing the first Context-Aware Anatomical Routing Attention (CARA) mechanism, which achieves precise anatomical guidance by dynamically selecting anatomical priors and modulating decoder attention intensity. Methodologically, the framework integrates anatomy-grounded pretraining with task-specific attention strategies. Experimental results demonstrate that the proposed model significantly improves regional localization accuracy and clinical evaluation performance while exhibiting strong generalization capability across external cohorts.
📝 Abstract
Cardiac magnetic resonance imaging (CMR) enables assessment of cardiac anatomy, ventricular function, and myocardial tissue characteristics. Clinicians interpret these images by identifying cardiac structures and focusing on the regions relevant to each clinical question, motivating anatomically guided vision-language models (VLMs). Yet CMR-specific supervision for anatomical localisation and clinical question answering remains limited. To address this gap, we investigate fine-grained CMR visual question answering through anatomical grounding and guided attention. We construct 128,915 anatomical-grounding and 42,799 clinical QA pairs across short-axis cine, late gadolinium enhancement, and long-axis cine. These datasets support anatomical recognition, localisation, and clinical assessment without requiring paired reports for individual training images. To help the model learn where to look, we introduce Cardiac Anatomy-Routed Attention (CARA), which selects predicted anatomical priors according to the question and guides decoder attention with learned task-specific strengths. Combining anatomical grounding pretraining with CARA yields our model, CARA-VL. Experiments demonstrate CARA-VL's strengths in clinical assessment and regional localisation across CMR imaging settings, with promising generalization to an external clinical cohort. Together, our data and method provide a practical framework for studying and advancing cardiac visual understanding in VLMs. We will release the QA data derived from public datasets upon publication.
Problem

Research questions and friction points this paper is trying to address.

Cardiac MRI
Vision-Language Models
Visual Question Answering
Anatomical Grounding
Guided Attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Anatomical Grounding
Guided Attention
Cardiac MRI
Visual Question Answering
💼 Related Jobs
No related jobs found.
Bangwei Guo
Bangwei Guo
Rutgers University
Medical imagecomputer vision
Xiao Chen
Xiao Chen
United Imaging Intelligence
MRIFast ImagingCompressed SensingArtificial Intelligence
B
Boris Mailhe
United Imaging Intelligence, Boston, MA, USA
J
Jia Yao
University of Texas Southwestern Medical Center, TX, USA
Y
Yiqing Wang
Duke University, NC, USA
A
Ankush Mukherjee
United Imaging Intelligence, Boston, MA, USA
Yikang Liu
Yikang Liu
Shanghai Jiao Tong University
Computational Linguistics
Z
Zheyuan Zhang
United Imaging Intelligence, Boston, MA, USA
H
Hang Yu
United Imaging Intelligence, Boston, MA, USA
Terrence Chen
Terrence Chen
UII America, Inc.
Medical ImagingImage-guided Interventions and SurgeryArtificial IntelligenceComputer Vision
Shanhui Sun
Shanhui Sun
UII America, Inc.
Machine LearningComputer VisionMedical imaging processingMedical Imaging and Virtual Reality