🤖 AI Summary
This study addresses the lack of multi-turn interactive and multi-label diagnostic benchmarks for evaluating clinical reasoning in multimorbidity scenarios. We propose a novel interactive evaluation framework grounded in clinical decision algorithms. By leveraging synthetic data generation and large language model (LLM)-simulated physician–patient dialogues, we construct a multi-label diagnostic benchmark supporting multi-turn inquiry alongside a theoretical reference model. Our analysis reveals a “single-hypothesis anchoring” behavioral pattern in LLMs during clinical reasoning. Experimental results demonstrate that even state-of-the-art models achieve less than 10% accuracy in complex comorbidity settings, and that multi-turn interaction paradoxically degrades their performance, further corroborating the presence of this cognitive bias.
📝 Abstract
Patients often have several co-occurring clinical conditions, and the findings needed to identify and disambiguate them emerge over the course of a consultation. Evaluating clinical reasoning in this setting requires both multi-turn interaction and multi-label diagnosis. We introduce CLIMB, a benchmark in which a doctor model interviews a simulated patient to recover a ground truth set of co-occurring clinical conditions. Cases are synthesized from clinical decision algorithms and diagnostic datasets, grounding multimorbid presentations in structured clinical knowledge. Across six frontier and open models, none recovers the exact set of conditions in more than 10% of interactive cases. Diagnostic performance declines when conditions co-occur, even when models receive the full clinical record and the true number of conditions. Interaction reduces performance further. In controlled experiments, models behave like single-hypothesis trackers: they anchor on the diagnosis suggested by the opening findings, keep questioning around it, and recover a second condition mainly when a finding in view points to it. Questioning them further does not complete the set but adds mostly wrong diagnoses. We formalise this pattern with a theoretical reference model of single-hypothesis tracking. The benchmark, generator, and evaluation code are available at https://anonymous.4open.science/r/CLIMB-8340.