InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses critical bottlenecks in the continual training of medical multimodal large language models, including data inefficiency, insufficient supervision signals, and poor output consistency, by proposing the InfiMed2 framework. Methodologically, we construct a 55.68-billion-token corpus and design a phase-aware data strategy to optimize pretraining. Furthermore, we introduce an answer-stability-based SFT reconstruction mechanism combined with evidence-focused mixing during learning rate decay, alongside RLVR to enhance fine-tuning quality. Ultimately, we release 4B and 27B general-purpose medical multimodal foundation models. The 4B model achieves an average score of 66.73%, surpassing Qwen3.5-9B, while the 27B model attains 73.72%, establishing new state-of-the-art performance among open-source counterparts.
📝 Abstract
Recent medical multimodal models have benefited from larger corpora, broader modality coverage, and stronger reasoning-oriented training, yet effective data design across continued pretraining (CPT) and post-training remains challenging. Medical sources vary substantially in structure, granularity, and information density, and their utility shifts as training progresses from broad knowledge acquisition to late-stage consolidation. Meanwhile, post-training is often dominated by short-form visual question answering, providing limited supervision for informative and answer-consistent explanations. We introduce InfiMed2, a family of 4B and 27B generalist medical multimodal foundation models built around stage-aware data design. We curate a 55.68B-token corpus that combines broad clinical knowledge with context-rich biomedical visual evidence through source-specific processing. Our CPT pipeline first adapts the vision encoder, then builds broad medical knowledge, and finally transitions to an evidence-focused data mixture during learning-rate decay. For supervised fine-tuning (SFT), we regenerate visual question-answering responses using answer stability, answer-masked reconstruction, and correctness-constrained selection to produce more informative and answer-consistent supervision. The 4B model is further optimized with reinforcement learning with verifiable rewards (RLVR). Across five medical multimodal benchmarks, InfiMed2-4B achieves 66.73% mean accuracy after RLVR, surpassing the larger Qwen3.5-9B, while InfiMed2-27B reaches 73.72%, the highest among the evaluated open-weight models.
Problem

Research questions and friction points this paper is trying to address.

medical multimodal models
data design
continued pretraining
post-training
visual question answering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Foundation Model
Stage-aware Data Design
Stability-Aware Supervision
Continued Pretraining
RLVR
🔎 Similar Papers
No similar papers found.
G
Guanghao Zhu
The Hong Kong Polytechnic University
Z
Zeyu Liu
The Hong Kong Polytechnic University
Z
Zhitian Hou
The Hong Kong Polytechnic University
Pengkai Wang
Pengkai Wang
Zhejiang University | HK Polytechnic University
control & optimizationAI for ScienceLLM-post training
Zhijie Sang
Zhijie Sang
Microsoft
NLP
S
Shuo Cai
The Hong Kong Polytechnic University
Y
Yang Yu
The Hong Kong Polytechnic University
Y
Yuanyi Wang
The Hong Kong Polytechnic University
Yanggan Gu
Yanggan Gu
Soochow University
Natural Language ProcessingLanguage Model
Congkai Xie
Congkai Xie
Reallm Labs
J
Jianmin Wu
The Hong Kong Polytechnic University, PolyU-Daya Bay Technology and Innovation Research Institute
Hongxia Yang
Hongxia Yang
Professor, HK Polytechnic University
Machine LearningGenerative AICognitive IntelligenceStatistical Modeling