🤖 AI Summary
This work addresses a critical gap in existing medical vision-language model benchmarks, which predominantly focus on disease classification or report generation while neglecting systematic evaluation of radiological spatial and anatomical reasoning capabilities. To bridge this gap, the authors introduce SPARC-Rad—the first high-quality multimodal benchmark specifically designed to assess such reasoning skills. It encompasses three major imaging modalities (CT, MRI, and X-ray) across five anatomical regions and includes 300 expert-annotated image-question pairs. The benchmark is accompanied by a standardized evaluation pipeline integrating prompt engineering, response normalization, LLM-as-judge automated scoring, human review, and binary correctness adjudication. This framework enables fine-grained subgroup analysis, model failure diagnosis, and pre-deployment validation, offering a reproducible and extensible foundation for evaluating medical vision-language models.
📝 Abstract
Vision-language models (VLMs) are increasingly being evaluated for medical imaging, but many available benchmarks emphasize disease classification, report generation, or broad visual question answering rather than the spatial and anatomical reasoning required for radiology. We developed the Spatial Perception and Anatomical Reasoning in Clinical Radiology (SPARC-Rad) Benchmark, a manually curated multimodal benchmark dataset and evaluation pipeline for assessing these capabilities in radiology VLMs. SPARC-Rad includes 300 image-question pairs derived from healthy control imaging studies in The Cancer Imaging Archive (TCIA), spanning CT, MRI, and radiography across the abdomen, chest, breast, neuro, and musculoskeletal categories. Radiology trainees manually designed and annotated questions to evaluate anatomical identification, localization, laterality, regional recognition, device identification, and inter-structure spatial relationships. The evaluation pipeline supports standardized prompting, structured output collection, response normalization, LLM-as-judge grading, human quality review, binary correctness scoring, and subgroup analysis by modality, anatomy, and reasoning type. SPARC-Rad provides a reusable framework for evaluating whether VLMs can provide reasoning for radiologic anatomy as a spatial system, supporting future model development, failure-mode analysis, and pre-deployment assessment.