OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of clinical recommendation misclassification caused by excessive refusal in large language models deployed within molecular tumor boards. To this end, we construct an open-source benchmark and a multimodal safety labeling system. Methodologically, we propose an auditing agent architecture that decouples verification from classification to mitigate label collapse, alongside a seven-module deterministic reasoning framework designed to precisely distinguish between evidence-supported recommendations and clinical warnings. Experimental results demonstrate that the proposed approach reduces the over-refusal rate to 6.7% while achieving a classification accuracy of 91.2%. This work provides an effective paradigm for enhancing the reliability and clinical utility of medical AI systems.
📝 Abstract
Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review because of incomplete information, poor ECOG performance status, or other clinical caveats. We introduce OpenMTB-Audit, an open-source benchmark of 500 synthetic non-small cell lung cancer cases spanning five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. Across eight large language model configurations, we identify pervasive over-refusal: all LLM configurations failed to retain the Partially Supported label in 83.3-100% of true Partially Supported cases, achieving high aggregate safety scores through label collapse rather than clinically calibrated reasoning. To address this limitation, we developed MTB-AuditAgent, a deterministic seven-module framework separating evidence verification, missing-information detection, safety classification, and abstention. It reduces over-refusal to 6.7% and achieves 91.2% accuracy (95% CI: 88.6-93.6%). A two-oncologist annotation study found disagreement concentrated at the boundary between information sufficiency and treatment optimization, underscoring the need to preserve clinically meaningful distinctions.
Problem

Research questions and friction points this paper is trying to address.

Molecular Tumor Board
Over-Refusal
Large Language Models
Precision Oncology
Safety Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Over-Refusal
Molecular Tumor Board
Safety Evaluation
Deterministic Framework
Benchmark
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Negin Ashrafi
Negin Ashrafi
Graduate Student at University of Southern California
Machine LearningDeep LearningOptimizationNLPStatistical Learning
J
Jia Luo
Lowe Center for Thoracic Oncology, Dana-Farber Cancer Institute, and Harvard Medical School, Boston, MA, USA
S
Stacey M. Frumm
Lowe Center for Thoracic Oncology, Dana-Farber Cancer Institute, and Harvard Medical School, Boston, MA, USA
R
Roxana Daneshjou
Department of Biomedical Data Science, Stanford University, Stanford, CA, USA