🤖 AI Summary
This study addresses the limitation of existing benchmarks that evaluate perception or single-step reasoning in isolation, thereby failing to comprehensively assess multimodal understanding and dynamic decision-making capabilities in intraoperative anesthesia. We construct the first evaluation suite spanning the entire pipeline from perception through single- to multi-step decision-making. Leveraging expert-annotated data, we develop a domain-specific evaluator via supervised fine-tuning and preference alignment training, enabling fine-grained quantification of clinical correctness and safety. Experiments reveal significant deficiencies in mainstream multimodal large language models regarding visual grounding and safe decision-making, with the best-performing model achieving an mIoU of only 32.2 and a severe error rate of 17.5%. These findings demonstrate that conventional aggregate metrics are insufficient for ensuring longitudinal decision safety in clinical settings.
📝 Abstract
Intraoperative anesthesia requires systems to interpret evolving multimodal evidence, recommend timely management, and revise decisions as patient states change, yet existing benchmarks usually isolate perception or single-point reasoning. We introduce AnesTRACE, an evaluation suite comprising AnesTRACE-Bench and AnesTRACE-Eval. Built from public perioperative datasets with anesthesiologist annotation, AnesTRACE-Bench evaluates Intraoperative Perception, Single-point Anesthesia Decision-Making, and Multi-step Anesthesia Decision-Making. AnesTRACE-Eval assesses open-ended responses through anesthesiologist-defined criteria for Clinical Correctness, Evidence Grounding, Task Completeness, and Safety, with Temporal Consistency for multi-step decisions; its domain-specific evaluator is trained by supervised fine-tuning and preference alignment on expert-reviewed judgments. Across more than 30 models, fine-grained visual grounding and intervention selection remain difficult: the leading model reaches only 32.2 mIoU for TEE visual grounding and retains a 17.5\% Major/Critical Safety Error Rate in multi-step management. Evaluator alignment with anesthesiologists improves across both training stages, while the best decision quality is accompanied by a 74.3-second P95 Latency. These results show that aggregate performance alone does not establish safe, timely longitudinal decision-making. We release our code at https://github.com/zjuDBxAI/AnesTRACE.