CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the difficulty of existing vision-language models (VLMs) in reliably extracting quantitative metrics from multidimensional cardiac MRI by proposing a tool-augmented VLM. The method leverages an interleaved reasoning mechanism to invoke a medical image analysis toolchain, including segmentation and volumetry, for quantitative assessment. A core innovation is the introduction of a Group Relative Policy Optimization (GRPO) training strategy with a conditional tool-use reward, alongside the construction of a multi-cohort benchmark to enhance tool-calling reliability. Through joint optimization via supervised fine-tuning and GRPO, the model achieves a Pass@1 of 35.9% and a tool-calling accuracy of 99.8%, improving ventricular measurement precision by over 20% compared to direct prediction.
📝 Abstract
Cardiovascular magnetic resonance (CMR), including cine imaging, is a reference standard for the noninvasive assessment of cardiac morphology and ventricular function. Cine CMR interpretation integrates qualitative visual assessment with quantitative measurements of ventricular volumes, ejection fraction, myocardial mass, wall thickness, and regional wall motion. Current medical vision-language models (VLMs) cannot reliably derive quantitative measurements from multidimensional cine images without analysis tools. We present CineMR, a tool-augmented VLM that invokes cardiac image-analysis tools and integrates their outputs into interleaved reasoning for quantitative CMR assessment. We also construct a multi-cohort visual question answering benchmark covering quantitative metric extraction, multiclass diagnosis, and differential diagnosis, together with tools for segmentation, phase selection, volumetry, morphometry, and regional wall motion analysis. CineMR is trained with supervised fine-tuning (SFT) on tool-interaction traces followed by Group Relative Policy Optimization (GRPO) with conditional tool-use rewards. On the multi-cohort cine CMR benchmark, CineMR achieves 35.9% pass@1 and 58.9% pass@4, compared with 1.5% pass@1 for the Qwen3-VL-8B backbone and 0.0% and 7.0% pass@1 for LLaVA-Med v1.5 and MedGemma-4B, respectively. Correct tool invocation reaches 99.8% after GRPO, up from 78.9% after SFT. Live tool outputs improve ventricular measurement accuracy by 20.4--23.7% over direct model predictions, and removing all tools reduces pass@1 from 35.9% to 27.9%. These results highlight the importance of reliable tool use for quantitative cine CMR reasoning and support CineMR as a promising approach for assistive cardiac image assessment. Code, benchmark resources, and model weights are available at https://github.com/AI-MIND-Lab/CineMR.
Problem

Research questions and friction points this paper is trying to address.

Cardiovascular Magnetic Resonance
Vision-Language Models
Quantitative Assessment
Cine Imaging
Tool Integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tool-Integrated VLM
Cardiac MRI
Group Relative Policy Optimization (GRPO)
Visual Question Answering
Quantitative Reasoning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kunyang Li
Institute for Artificial Intelligence, University of Central Florida
Hai Nguyen
Hai Nguyen
Applied Scientist, Amazon Robotics | Ph.D. Northeastern University
RoboticsReinforcement Learning
J
Joshua Lowe
Institute for Artificial Intelligence, University of Central Florida
Chenguang Zhao
Chenguang Zhao
Nemours Cardiac Center, Nemours Children’s Hospital, Florida
P
Peace C. Madueme
Nemours Cardiac Center, Nemours Children’s Hospital, Florida
M
Mehdi Hedjazi Moghari
Children’s Heart Center, WVU Golisano Children’s, West Virginia
Mubarak Shah
Mubarak Shah
Trustee Chair Professor of Computer Science, University of Central Florida
Computer Vision
Pegah Khosravi
Pegah Khosravi
Associate Professor, Institute for Artificial Intelligence, University of Central Florida
Artificial IntelligenceComputer VisionMedical Image AnalysisMultimodal Learning
Y
Yuzhang Zhang
Institute for Artificial Intelligence, University of Central Florida