π€ AI Summary
This study addresses the critical gap in current vision-language models for lumbar MRI report generation, which often produce fluent yet clinically inaccurate diagnoses that standard natural language metrics fail to capture. To bridge this divide, the authors introduce the first clinical semantic evaluation benchmark tailored to lumbar MRI and propose an architecture-agnostic enhancement framework. This approach leverages a semi-supervised U-Net++ to generate intervertebral discβlevel abnormality heatmaps, providing spatially explicit guidance to the vision-language model and thereby enhancing its anatomical sensitivity and diagnostic reliability. The method significantly improves clinical correctness while offering interpretable visual evidence, revealing for the first time the disconnect between conventional language metrics and clinical accuracy, and advancing vision-language models toward real-world clinical utility.
π Abstract
Radiology reporting is time-consuming and subject to inter-rater variability, making automated report generation an attractive clinical application for Vision-Language Models (VLMs). We benchmark state-of-the-art VLMs on lumbar spine MRI with a focus on diagnostic accuracy and demonstrate that standard lexical and semantic metrics poorly reflect clinical correctness: fluent, well-structured reports can score highly while containing clinically meaningful diagnostic errors. To address this failure mode, we propose an architecture-agnostic framework that augments VLM inputs with spatially localized, disc-level anomaly heatmaps generated by a semi-supervised U-Net++ model. These heatmaps both improve anatomical sensitivity through explicit visual grounding and provide an independent interpretability output for clinical oversight, moving us closer to diagnostically reliable, visually grounded VLMs for lumbar spine MRI interpretation.