OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of existing benchmarks to capture authentic multidisciplinary tumor board (MDT) discussion trajectories by constructing the first large-scale MDT dialogue dataset derived from publicly available video transcripts, comprising 611 patients and 19,000 discussion turns. A dual-mode evaluation framework is designed to quantify clinical equivalence and consensus alignment. Methodologically, an automated data curation pipeline is combined with supervised fine-tuning and reinforcement learning techniques to optimize model performance, with rigorous medical expert review ensuring data quality. The findings reveal critical limitations of state-of-the-art large language models in complex medical decision-making while validating the potential of leveraging authentic discussion trajectories to enhance the clinical adaptability of such models.
📝 Abstract
Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchmark with 611 patient cases and 19,157 discussion turns across ten specialist roles, transcribed from 12,534 minutes of publicly available tumor board recordings on YouTube. The benchmark evaluates two settings: SPECIALIST TURN, in which an LLM responds to a clinically significant question posed during a real discussion, and BOARD SIMULATION, in which it generates an entire back-and-forth discussion and reaches a consensus on therapy recommendations, surgical plans, next actions and clinical trial matching. Evaluation of 14 general-purpose frontier and medical LLMs reveals substantial limitations: the best models score 3.43 out of 5 in clinical equivalence to specialist answers and 2.78 out of 5 in alignment with recorded board conclusions. Supervised finetuning and reinforcement learning improve performance on a held-out test set, suggesting that real-world discussion trajectories can support model adaptation. Three M.D. experts review a subset of the benchmark, finding high information coverage and factuality of patient cases and strong fidelity of extracted consensus conclusions. We will release OpenTumorBoard and its automated curation pipeline to support the development and evaluation of LLMs for multidisciplinary, personalized cancer decision-making.
Problem

Research questions and friction points this paper is trying to address.

Multidisciplinary Tumor Board
Benchmark
Large Language Models
Clinical Decision-Making
Oncology
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multidisciplinary Tumor Board
Benchmark
Large Language Models
Board Simulation
Automated Curation Pipeline
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Anqi Li
Rice University
Z
Zhixuan Ge
Rice University
Y
Yixuan Duan
Rice University
J
Jiarong Qian
Rice University
C
Chi-Yu Chen
University of Illinois at Urbana-Champaign
M
MingYu Lu
University of Washington
H
Huan-Yu Hsu
National Yang Ming Chiao Tung University
Y
Yu Gu
Microsoft
Yue Guo
Yue Guo
Assistant Professor, University of Illinois Urbana-Champaign
NLPmedicine
Sheng Wang
Sheng Wang
Assistant Professor at University of Washington
machine learningcomputational biologycancer genomicsdrug discovery
W
Wei Qiu
Rice University
Hanwen Xu
Hanwen Xu
University of Washington
Artificial IntelligencePrecision Health