Beyond Report Imitation: Clinically Aware Multi-Image Ultrasound Report Generation from Visible Evidence

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misalignment between visual evidence and clinical supervision caused by raw report imitation in multi-view ultrasound report generation. To this end, we propose CAMEO, a three-stage framework comprising visual-language primitive learning, cross-view evidence grounding, and clinically aware preference alignment. By constructing evidence-grounded supervision alongside a clinical-error-guided preference optimization mechanism, and by integrating cross-view evidence distillation with multi-image question answering, CAMEO effectively overcomes the limitations of conservative templating and diagnostic inversion. Evaluated on the USReport benchmark, the proposed method achieves a BLEU-1 score of 0.40, a ROUGE-1 score of 0.45, and elevates the ClinicalScore to 74.20, demonstrating a substantial improvement in the clinical reliability of generated reports.
📝 Abstract
Generating ultrasound reports from multiple images requires aggregating clinical evidence across views, yet archived key frames capture only part of the dynamic examination. Raw-report imitation is therefore misaligned with visual supervision: content that is clinically valid for the full examination may be unverifiable from the images available to a model. This gap creates a clinical behavior alignment problem. A model must preserve visible findings, avoid diagnostic reversals and unsupported completion, and not collapse into conservative templates. We propose CAMEO, a Clinically Aware Multi-image Evidence-grounded Orchestration framework for ultrasound report generation. Stage I learns ultrasound visual-language primitives; Stage II performs Cross-View Evidence Grounding by distilling trusted visible report points into multi-image QA and report-style supervision; and Stage III performs Clinically Aware Preference Alignment using clinical-error-oriented preference pairs. From USReport, we construct USReport-Distilled with 17,670 evidence-grounded paired-image training instances and USReport-Pref with 21,869 preference pairs; we additionally use 25,631 PubMedVision-US ultrasound instruction samples for domain adaptation and multi-image instruction tuning. On the primary USReport-Distilled benchmark, CAMEO improves over EchoVLM from 0.25 to 0.40 BLEU-1, 0.28 to 0.45 ROUGE-1, and 0.27 to 0.43 METEOR, while raising ClinicalScore from 55.02 to 74.20. These results underscore the value of evidence-grounded supervision, clinically aware alignment, and clinically structured evaluation for reliable ultrasound report generation.
Problem

Research questions and friction points this paper is trying to address.

ultrasound report generation
multi-image evidence grounding
clinical behavior alignment
visible evidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-image Ultrasound Report Generation
Cross-View Evidence Grounding
Clinically Aware Preference Alignment
Evidence-grounded Supervision
Report Distillation
💼 Related Jobs
No related jobs found.
Y
Yuchen Yang
State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China
Xin Wang
Xin Wang
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
Biomedical Engineering
L
Lufan Wang
State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China
Y
Yinghong Pan
State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China
Y
Yujuan Feng
College of Computer Science, Beijing University of Technology, Beijing, China
Yuqing Yang
Yuqing Yang
Beijing University of Posts and Telecommunications
Machine LearningBioinformaticsMedical Informatics