RadSight: Towards Perceptually Reliable Multimodal Radiology Image Understanding

📅 2026-07-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the unreliable performance of existing medical multimodal large language models in assessing fundamental lesion attributes—such as location, size, and density—due to their limited visual perception capabilities. To this end, we introduce Perception-Bench, a large-scale evaluation benchmark comprising 1.13 million samples, and propose RadSight, a novel model featuring a dual-path 2D/3D vision encoder that natively preserves spatial structure. RadSight decomposes medical image understanding into four stages: vision–language alignment, fine-grained perception, clinical diagnosis, and explanation generation, trained via progressive curriculum learning on 8.37 million perception-oriented samples. Experiments demonstrate that RadSight consistently outperforms prior methods across all six dimensions of Perception-Bench, with particularly notable gains in spatial localization and diagnostic accuracy, while also achieving consistent performance improvements on multiple public 2D and 3D medical imaging benchmarks.
📝 Abstract
Medical multimodal large language models (MLLMs) are increasingly expected to perform complex image understanding tasks, yet their reliability is often compromised by frequent errors in visual interpretation. To systematically trace these failures, we traverse the hierarchy from high-level clinical tasks down to fundamental visual perception. We therefore introduce Perception-Bench, a large-scale benchmark comprising 1.13 million samples that assesses medical MLLMs across six dimensions: attribute judgment, spatial grounding, spatial understanding, disease prediction, anomaly detection, and report generation, spanning both 2D and 3D radiology images. Our analysis on Perception-Bench reveals that existing MLLMs lack the ability to capture even the most basic lesion attributes, such as location, size, and density. This inability to ground clinical outputs in primary visual evidence reveals that the models' diagnostic unreliability is rooted in a critical but overlooked bottleneck in low-level visual perception. Motivated by this, we propose RadSight, a perception-driven MLLM built upon a dual 2D/3D encoder architecture that preserves native imaging spatial structures. RadSight formulates medical image understanding as a four-stage progressive process: visual-language alignment, fine-grained visual perception, clinical diagnosis, and diagnostic interpretation. The model is trained on an 8.37 million perception-oriented corpus using progressive curriculum learning. On Perception-Bench, RadSight consistently outperforms existing MLLMs across all six evaluation dimensions, with particularly strong gains in spatial grounding and clinical diagnosis. It also achieves consistent improvements on public 2D and 3D medical benchmarks, further demonstrating that robust low-level visual perception is a critical foundation for reliable clinical understanding. Code and model will be publicly available.
Problem

Research questions and friction points this paper is trying to address.

medical multimodal large language models
visual perception
radiology image understanding
diagnostic reliability
lesion attributes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Perception-Bench
RadSight
multimodal medical LLM
visual perception
2D/3D radiology understanding