🤖 AI Summary
This study addresses the limitations of existing benchmarks in evaluating clinical reasoning over longitudinal, multimodal, and volumetric evidence chains. We construct a visual question answering (VQA) benchmark comprising 7,273 instances derived from brain MRI reports, leveraging approximately 29,000 registered 3D MRI scans. To our knowledge, this work introduces the first framework for longitudinal multimodal volumetric evidence tracing, incorporating a six-step decomposed auditing mechanism and a five-tier reasoning hierarchy evaluated through standard vision-language model (VLM) interfaces. Experiments across 20 VLM configurations reveal that while current models can identify isolated cues, they rarely generate grounded longitudinal explanations. By bridging the gap left by single-image evaluations, this research exposes fundamental compositional reasoning bottlenecks in contemporary VLMs.
📝 Abstract
Brain MRI interpretation is a longitudinal clinical reasoning problem: radiologists compare serial studies, integrate information across MRI sequences, localize findings within volumetric anatomy, and translate this evidence into report-grounded assessments. Existing medical VQA and 3D imaging benchmarks capture important parts of this workflow, but often evaluate brain MRI through isolated images, static volumes, or ungrounded report-style answers, thereby obscuring failures in the evidence chain that support clinical validity. We introduce BrainTRACE, a report-grounded benchmark for evaluating whether vision-language models can trace the evidence structure required for longitudinal brain MRI interpretation. BrainTRACE contains 7,273 scored VQA instances derived from 1,778 longitudinal patients, 7,299 MRI studies, and approximately 29k co-registered 3D MRI sequence volumes. The benchmark is organized by five levels of clinical reasoning, from acquisition recognition to case-level synthesis, and by evidence demands covering longitudinal comparison, report-grounded references, multi-sequence integration, and volumetric spatial evidence. BrainTRACE supports rendered inputs compatible with standard VLM interfaces, a 3D-evidence condition, and a decomposed case-reasoning track that audits six steps in a longitudinal evidence chain. Evaluation of 20 VLM configurations shows that current systems can identify isolated visual cues but rarely compose them into grounded longitudinal interpretations. We release the benchmark specification, evaluation lists, scoring implementation, scoring rubrics, and audit-record format to support reproducible progress in brain MRI VLM evaluation.