MedKIT: Evaluating Knowledge Integration and Generalization in Large Language Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that medical large language models struggle to effectively apply updated knowledge in complex clinical scenarios. To this end, we propose MedKIT, a benchmark that transcends traditional evaluations limited to factual recall by introducing a multidimensional assessment framework encompassing relational, compositional, and operational transfer capabilities. This framework systematically examines model integration and generalization under sequences of clinical knowledge updates. Through a large-scale empirical study evaluating twelve updating strategies across five model architectures, our work reveals a significant gap between "recallable" and "usable" knowledge. The findings demonstrate that existing methods remain insufficient for complex clinical reasoning tasks. Ultimately, MedKIT establishes a new paradigm for research on knowledge updating in medical foundation models.
📝 Abstract
Constantly evolving real-world knowledge necessitates models to be updated continuously. Especially in medicine, as clinical evidence changes over time, outdated knowledge can pose safety risks. Existing evaluations of knowledge integration focus on factual recall, offering limited insight into whether newly integrated knowledge is actually usable. Our benchmark MedKIT (Medical Knowledge Integration and Transfer) provides a granular evaluation of how models integrate and apply knowledge under realistic sequences of clinical updates. Each instance corresponds to a factual update derived from clinical evidence, paired with targeted probes that assess transfer across lexical variation, relational transformations, compositional reasoning, and open-ended operationalization, as well as locality tests for knowledge preservation. Using MedKIT, we conduct a large-scale empirical study of 12 knowledge integration strategies across 5 diverse models, including both general-purpose and medical LLMs. Our results reveal a consistent gap between recall and usable knowledge: while most methods achieve strong gains on the original update task and under lexical variation, relational generalization is limited, and no method yields meaningful improvements on compositional or operational tasks. These findings highlight a fundamental challenge in knowledge integration and position MedKIT as a testbed for developing methods that make newly integrated knowledge more consistently usable across tasks and contexts.
Problem

Research questions and friction points this paper is trying to address.

Knowledge Integration
Large Language Models
Medical Knowledge
Generalization
Evaluation Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Knowledge Integration
Large Language Models
Medical Benchmark
Generalization
Compositional Reasoning
🔎 Similar Papers