UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of pixel-level evidence alignment and fine-grained localization capabilities in large vision-language models (VLMs) for ultrasound imaging. To this end, we construct a multi-task benchmark encompassing 40 datasets to systematically evaluate VLM localization performance across segmentation, visual question answering, and report generation tasks. Furthermore, we propose UltraG-Agent, a framework that deeply integrates the semantic reasoning capabilities of VLMs with ultrasound-specific segmentation models. Our investigation reveals a significant discrepancy between semantic understanding and pixel-level localization. Experimental results demonstrate that the proposed approach effectively enhances prediction accuracy and visual grounding performance, thereby establishing a novel paradigm for multimodal ultrasound analysis.
📝 Abstract
Ultrasound is one of the most widely used medical imaging modalities, and recent large vision-language models(VLMs) have shown increasing capabilities in ultrasound image understanding. However, these models fail to provide pixel-level visual evidence aligned with their semantic predictions, and their fine-grained grounding capability in ultrasound remains largely unclear. We introduce UltraG-Bench, a large-scale multi-task benchmark for evaluating pixel-level evidence grounding in ultrasound. UltraG-Bench is built by annotating 40 public ultrasound segmentation datasets spanning 13 anatomical categories, and comprises three progressive tasks: instruction-guided segmentation, evidence-grounded VQA, and evidence-grounded report generation, with 331125, 666779, and 138832 annotations, respectively. Comprehensive evaluation of 14 state-of-the-art models reveals a substantial gap between semantic understanding and fine-grained pixel-level localization. We further propose UltraG-Agent, which combines the semantic reasoning capabilities of a VLM with the ultrasound-specific segmentation capability of UltraSAM3. Experiments show that UltraG-Agent substantially improves both semantic prediction and pixel-level visual grounding. Our dataset and code are available at https://github.com/zhuqh19/UltraG-Bench.
Problem

Research questions and friction points this paper is trying to address.

Ultrasound
Large Vision-Language Models
Pixel-level Evidence Grounding
Visual Grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ultrasound
Vision-Language Models
Pixel-level Grounding
Multi-task Benchmark
Segmentation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Q
Quanhao Zhu
Dalian University of Technology
Bo Xu
Bo Xu
Dalian University of Technology
Natural Language ProcessingInformation RetrievalMedical DialoguePsychological Computing
R
Rui Lin
Dalian University of Technology
C
Chenyuan Wang
Dalian University of Technology
Y
Yu Shao
Dalian University of Technology
B
Boling Zhu
Dalian University of Technology
J
Jiuyan Sun
Dalian University of Technology
L
Liang Zhao
Dalian University of Technology
Hongfei Lin
Hongfei Lin
DalianUniversity of Technology
natural language processing,sentimental analysistext miningsocial computing
Feng Xia
Feng Xia
Professor, School of Computing Technologies, RMIT University
Artificial IntelligenceGraph LearningBrainRoboticsCyber-Physical Systems