SPARC-Rad: A Multimodal Benchmark Dataset and Evaluation Pipeline for Spatial and Anatomical Reasoning in Radiology Vision-Language Models

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical gap in existing medical vision-language model benchmarks, which predominantly focus on disease classification or report generation while neglecting systematic evaluation of radiological spatial and anatomical reasoning capabilities. To bridge this gap, the authors introduce SPARC-Rad—the first high-quality multimodal benchmark specifically designed to assess such reasoning skills. It encompasses three major imaging modalities (CT, MRI, and X-ray) across five anatomical regions and includes 300 expert-annotated image-question pairs. The benchmark is accompanied by a standardized evaluation pipeline integrating prompt engineering, response normalization, LLM-as-judge automated scoring, human review, and binary correctness adjudication. This framework enables fine-grained subgroup analysis, model failure diagnosis, and pre-deployment validation, offering a reproducible and extensible foundation for evaluating medical vision-language models.
📝 Abstract
Vision-language models (VLMs) are increasingly being evaluated for medical imaging, but many available benchmarks emphasize disease classification, report generation, or broad visual question answering rather than the spatial and anatomical reasoning required for radiology. We developed the Spatial Perception and Anatomical Reasoning in Clinical Radiology (SPARC-Rad) Benchmark, a manually curated multimodal benchmark dataset and evaluation pipeline for assessing these capabilities in radiology VLMs. SPARC-Rad includes 300 image-question pairs derived from healthy control imaging studies in The Cancer Imaging Archive (TCIA), spanning CT, MRI, and radiography across the abdomen, chest, breast, neuro, and musculoskeletal categories. Radiology trainees manually designed and annotated questions to evaluate anatomical identification, localization, laterality, regional recognition, device identification, and inter-structure spatial relationships. The evaluation pipeline supports standardized prompting, structured output collection, response normalization, LLM-as-judge grading, human quality review, binary correctness scoring, and subgroup analysis by modality, anatomy, and reasoning type. SPARC-Rad provides a reusable framework for evaluating whether VLMs can provide reasoning for radiologic anatomy as a spatial system, supporting future model development, failure-mode analysis, and pre-deployment assessment.
Problem

Research questions and friction points this paper is trying to address.

spatial reasoning
anatomical reasoning
radiology
vision-language models
medical imaging benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

spatial reasoning
anatomical reasoning
vision-language models
radiology benchmark
multimodal evaluation
S
Satvik Tripathi
Department of Radiology, Perelman School of Medicine, University of Pennsylvania, Philadelphia, Pennsylvania, USA
M
Mustafa Ege Seker
Department of Radiology, University of Wisconsin–Madison School of Medicine and Public Health, Madison, Wisconsin, USA
K
Kristian Quevada
Department of Radiology, Perelman School of Medicine, University of Pennsylvania, Philadelphia, Pennsylvania, USA; Department of Radiology, Cooper University Hospital, Cooper Medical School of Rowan University, Camden, New Jersey, USA
E
Ebubechukwu D Enwerem
Department of Radiology, Perelman School of Medicine, University of Pennsylvania, Philadelphia, Pennsylvania, USA; College of Computing and Informatics, Drexel University, Philadelphia, Pennsylvania, USA
P
Pratham Khandelwal
Department of Computer Science and Engineering, University of Minnesota Twin Cities, Minneapolis, Minnesota, USA
E
Emine Meltem
Istanbul Training and Research Hospital, Department of Radiology, İstanbul, Türkiye
B
Bera Koca
Department of Radiology, School of Medicine, Acıbadem Mehmet Ali Aydınlar University, Istanbul, Türkiye
Shahriar Faghani
Shahriar Faghani
Adjunct Assistant Professor, Department of Radiology, Mayo Clinic, MN, USA
RadiologyNeuroradiologyDeep LearningImaging InformaticsUncertainty Quantification
J
Jacinta Arnold
Department of Radiology, Perelman School of Medicine, University of Pennsylvania, Philadelphia, Pennsylvania, USA; UC Davis Graduate School of Management, Davis, CA
D
Dania Daye
Department of Radiology, University of Wisconsin–Madison School of Medicine and Public Health, Madison, Wisconsin, USA
T
Tessa S. Cook
Department of Radiology, Perelman School of Medicine, University of Pennsylvania, Philadelphia, Pennsylvania, USA