🤖 AI Summary
This study addresses the limitations of large language models in medical reasoning, specifically their reliance on static knowledge and the prohibitive costs of expert supervision. To overcome these challenges, we propose a self-evolving framework tailored for open-domain medical applications, wherein question-generation and problem-solving agents collaborate to achieve continuous optimization. Furthermore, this work introduces a novel controlled knowledge accumulation strategy that integrates multi-turn evidence-guided reasoning, external tool invocation, and both exploratory and persistent knowledge management techniques to ensure iterative reliability. This mechanism effectively transcends the static knowledge bottleneck inherent in existing approaches. Extensive evaluations across five benchmarks demonstrate that the proposed framework significantly outperforms current baselines, achieving an average accuracy improvement of up to 13.7 percentage points.
📝 Abstract
Large language models (LLMs) have shown promise in medical question answering and clinical reasoning, yet their improvement remains constrained by static parametric knowledge and costly expert supervision. Self-evolving agents offer a promising alternative by enabling models to improve through iterative task generation and problem-solving. However, most existing self-evolving methods are designed for easily verifiable domains such as mathematics and coding, where solutions can be checked by exact answers or executable programs. Medical reasoning is fundamentally different: it is open-ended, knowledge-intensive, and often only partially verifiable. We present MedZERO, a self-evolving framework for open-ended medical reasoning. MedZERO couples an Examiner that generates frontier medical question-option pairs with a Reasoner that solves them through evidence-grounded multi-turn reasoning with external knowledge tools. To support reliable, continual improvement, MedZERO adopts controlled knowledge accumulation, which maintains temporary exploratory knowledge and curated persistent knowledge in reasoning. We evaluate MedZERO on five public medical reasoning benchmarks using 4B- and 8B-scale base models under open-ended evaluation. Across all settings, MedZERO consistently outperforms the underlying base models and prior self-evolving baselines, achieving up to 13.7 average accuracy-point gains over the next-best self-evolving baseline.