ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses key vision-centric challenges in deploying multimodal large models for clinical settings, particularly insufficient understanding of heterogeneous 2D/3D medical images and the absence of fine-grained, clinically aligned evaluation protocols. The authors propose a cascaded spatial-aware local fusion encoder to unify native 2D and 3D medical image modeling and introduce a Vision-Grounded evaluation framework—comprising MedIF-Bench and region-of-interest (RoI)-grounded metrics—that enables, for the first time, automatic, fact-based, and clinically aligned assessment grounded in anatomical regions. Integrated with retrieval augmentation and agent-based tool invocation, the model outperforms existing open-source counterparts on 20 out of 24 benchmarks and surpasses GPT-5.2 and Gemini-3-Flash on 13 out of 16 tasks. Blind evaluations by radiologists confirm its superior report quality, with RoI-grounded metrics showing the strongest correlation with expert judgments.
📝 Abstract
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understanding within a fused encoder. We further introduce a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for clinically aligned and factualness-driven report generation evaluation. We show that ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks---spanning visual question answering, report generation, and instruction following---as well as textual medical tasks, outperforming leading open-source medical MLLMs (\textit{e.g.}, Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrating multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks, and can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirms that ClinFusion produces the highest-ranked reports, and validates our RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.
Problem

Research questions and friction points this paper is trying to address.

vision-centric
multimodal large language models
medical image understanding
clinical evaluation
factualness-driven assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-centric MLLM
Cascade Spatial-Aware Locality Fusion
3D medical image understanding
RoI-grounded evaluation
multimodal medical benchmarking
🔎 Similar Papers
2024-01-02IEEE International Conference on Bioinformatics and BiomedicineCitations: 0
Hangjie Yuan
Hangjie Yuan
Alibaba DAMO | ZJU | MMLab@NTU
Generative ModelsMultimodal ModelsFoundation ModelsVideo Understanding
Yichen Qian
Yichen Qian
Alibaba DAMO Academy
Computer VisionFace and GestureGenerative Adversarial Networks
Zhiwei Tang
Zhiwei Tang
Professor, University of Electronic Science and Technology of China
X
Xianzhe Xu
DAMO Academy, Alibaba Group, Beijing, China; Hupan Laboratory, Hangzhou, China
Lirong Wu
Lirong Wu
Zhejiang University & Westlake University
Geometric Deep LearningAI4Science
Sicheng Yang
Sicheng Yang
Tencent Robotics X
Robot
Jinwang Wang
Jinwang Wang
Researcher, Alibaba DAMO Academy
Computer VisionVisual Foundation Model
Pengju Wang
Pengju Wang
Chinese Academy of Science
Zhitao Zeng
Zhitao Zeng
National University of Singapore
Vision-Language Models
Yizeng Han
Yizeng Han
Alibaba DAMO Academy
Dynamic Neural NetworksEfficient Deep LearningComputer Vision
Y
Yan Xing
DAMO Academy, Alibaba Group, Hangzhou, China; Hupan Laboratory, Hangzhou, China
Shengxuan Luo
Shengxuan Luo
Center for Statistical Science and Department of Industrial Engineering, Tsinghua University
medical informaticsentity alignmenmachine translation
T
Tao Feng
Department of Computer Science and Technology, Tsinghua University, Beijing, China
Q
Qing Xie
Department of Radiology, The Affiliated Yangming Hospital of Ningbo University, Yuyao, China
W
Weigen Yao
Department of Radiology, The Affiliated Yangming Hospital of Ningbo University, Yuyao, China
Yi Yang
Yi Yang
Zhejiang University
multimediacomputer visionmachine learning
Zuozhu Liu
Zuozhu Liu
Assistant Professor, Zhejiang University/University of Illinois Urbana-Champaign
deep learningvision-language modelsmedical AI
J
Jiasheng Tang
DAMO Academy, Alibaba Group, Hangzhou, China; Hupan Laboratory, Hangzhou, China
S
Shaocheng Wang
Hepato-Pancreato-Biliary Center, Beijing Tsinghua Changgung Hospital, School of Clinical Medicine, Tsinghua Medicine, Tsinghua University, Beijing, China
Jitao Wang
Jitao Wang
University of Michigan
BiostatisticsMobile healthReinforcement learning
J
Jiahong Dong
Hepato-Pancreato-Biliary Center, Beijing Tsinghua Changgung Hospital, School of Clinical Medicine, Tsinghua Medicine, Tsinghua University, Beijing, China
Weihua Chen
Weihua Chen
Alibaba DAMO Academy, previously NLPR, CASIA
Computer Vision
Feng Xu
Feng Xu
Associate Professor of Tsinghua University
Computer Vision and Graphics
Fan Wang
Fan Wang
Alibaba DAMO Academy
Computer VisionMachine Learning