🤖 AI Summary
This study addresses the problem of clinical recommendation misclassification caused by excessive refusal in large language models deployed within molecular tumor boards. To this end, we construct an open-source benchmark and a multimodal safety labeling system. Methodologically, we propose an auditing agent architecture that decouples verification from classification to mitigate label collapse, alongside a seven-module deterministic reasoning framework designed to precisely distinguish between evidence-supported recommendations and clinical warnings. Experimental results demonstrate that the proposed approach reduces the over-refusal rate to 6.7% while achieving a classification accuracy of 91.2%. This work provides an effective paradigm for enhancing the reliability and clinical utility of medical AI systems.
📝 Abstract
Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review because of incomplete information, poor ECOG performance status, or other clinical caveats. We introduce OpenMTB-Audit, an open-source benchmark of 500 synthetic non-small cell lung cancer cases spanning five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. Across eight large language model configurations, we identify pervasive over-refusal: all LLM configurations failed to retain the Partially Supported label in 83.3-100% of true Partially Supported cases, achieving high aggregate safety scores through label collapse rather than clinically calibrated reasoning. To address this limitation, we developed MTB-AuditAgent, a deterministic seven-module framework separating evidence verification, missing-information detection, safety classification, and abstention. It reduces over-refusal to 6.7% and achieves 91.2% accuracy (95% CI: 88.6-93.6%). A two-oncologist annotation study found disagreement concentrated at the boundary between information sufficiency and treatment optimization, underscoring the need to preserve clinically meaningful distinctions.