🤖 AI Summary
This study addresses the unclear impact of agent frameworks on medical AI agent performance and their underlying mechanisms. To this end, it introduces MH-Lab, a benchmark with a controlled experimental environment that systematically evaluates five large language models across five distinct frameworks. It provides the first measurement of framework variance and decouples the contributions of individual mechanisms, such as context management and planning, through execution trajectory analysis. The findings reveal that frameworks and their interaction effects account for approximately 25% of the outcome variance, demonstrating that no universally optimal framework exists. This work offers quantitative evidence to guide framework selection and mechanism optimization for medical agents.
📝 Abstract
LLM agents are increasingly built for medical work and scored on clinical benchmarks. Each such score, however, comes from a model running inside an agent harness, the system that controls the loop between the model and its environment. An agent's score is therefore a property of a model--harness pair. For medical agents, how much outcomes change with the harness has rarely been measured. Measuring this change, and explaining it, raises two challenges. First, a harness comparison must change nothing but the harness and be repeated across models and kinds of task. Second, comparing whole harnesses leaves their mechanisms bundled together, so it cannot show when an individual mechanism helps. To address these challenges, we present MedicalHarness, a controlled study of models and agent harnesses on medical tasks. We first build MedicalHarnessBench to evaluate agents on $107$ tasks across four domains that each test a different harness capability. Using this benchmark, we run five open-weight models under five agent harnesses, changing only the harness within a comparison, and analyze both outcomes and execution traces. To study individual mechanisms, we build MH-Lab, a controlled harness that switches off context management, planning or tool exposure one at a time within a shared execution loop. We find that the harness and its interaction with the model account for about a quarter of the outcome variance, and that no single harness is best across models and tasks. Code and data are available at https://github.com/REAL-Lab-NU/MedicalHarness.