Representation-guided in-context learning for medical image interpretation with multimodal large language models

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决医疗图像解读中多模态大语言模型适应性问题,提出无需训练的表示引导上下文学习方法,通过检索对齐案例提高分类和视觉问答性能。
📝 Abstract
Medical image interpretation is central to diagnosis and care, yet adapting general-purpose multimodal large language models (MLLMs) often requires resource-intensive domain-specific fine-tuning. Here we introduce representation-guided in-context learning (RG-ICL), a training-free inference framework that retrieves query-aligned demonstrations using frozen encoders, without task-specific parameter updates. Across eight datasets spanning histopathology, radiology and retinal fundoscopy, RG-ICL improved classification (mean gain 20 percentage points) and visual question answering (VQA) (mean gain 13 percentage points) over no-context and conventional ICL, approaching or exceeding training-based comparators. Which cases were retrieved mattered more than how many: 6 query-aligned cases outperformed up to 32 randomly selected ones, whereas fixed or random cases often reduced accuracy below baseline. For VQA, aligning reference cases with both image content and question intent produced further gains. These findings indicate that for medical image interpretation, curating which reference cases an MLLM sees is a practical alternative to retraining it.
Problem

Research questions and friction points this paper is trying to address.

medical image interpretation
multimodal large language models
domain-specific fine-tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

representation-guided in-context learning
frozen encoders
query-aligned demonstrations
medical image interpretation
multimodal large language models
🔎 Similar Papers
M
Minda Zhao
Harvard AI and Robotics Lab, Schepens Eye Research Institute of Massachusetts Eye and Ear, Harvard Medical School, Boston, MA, USA
F
Fangyu Hu
Department of Ophthalmology, Massachusetts Eye and Ear, Harvard Medical School, Boston, MA, USA
Yan Luo
Yan Luo
Harvard University
Computer VisionMachine LearningBiomedical ImagingAI for Medicine
Yutong Yang
Yutong Yang
Mercedes-Benz AG R&D & University of Stuttgart
Computer VisionAutonomous Driving
J
Jiahui Cai
Harvard AI and Robotics Lab, Schepens Eye Research Institute of Massachusetts Eye and Ear, Harvard Medical School, Boston, MA, USA
K
Kaichen Zhou
Harvard AI and Robotics Lab, Schepens Eye Research Institute of Massachusetts Eye and Ear, Harvard Medical School, Boston, MA, USA
Manling Li
Manling Li
Assistant Professor at Northwestern University
Natural Language ProcessingVision-LanguageEmbodied Agents
Paul Liang
Paul Liang
Assistant Professor, Media Lab and EECS, MIT
machine learningartificial intelligencemultimodal interactionhuman-AI interaction
Yilun Du
Yilun Du
Harvard University
Artificial IntelligenceMachine LearningRoboticsComputer Vision
Lucy Q. Shen
Lucy Q. Shen
Associate Professor of Ophthalmology, Harvard Medical School
glaucomaimaging
Mengyu Wang
Mengyu Wang
Assistant Professor, Harvard Medical School
Artificial IntelligenceMachine LearningOphthalmologyGlaucomaComputational Mechanics