AnesTRACE: Benchmarking Intraoperative Anesthesia from Multimodal Perception to Multi-step Decision-Making

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing benchmarks that evaluate perception or single-step reasoning in isolation, thereby failing to comprehensively assess multimodal understanding and dynamic decision-making capabilities in intraoperative anesthesia. We construct the first evaluation suite spanning the entire pipeline from perception through single- to multi-step decision-making. Leveraging expert-annotated data, we develop a domain-specific evaluator via supervised fine-tuning and preference alignment training, enabling fine-grained quantification of clinical correctness and safety. Experiments reveal significant deficiencies in mainstream multimodal large language models regarding visual grounding and safe decision-making, with the best-performing model achieving an mIoU of only 32.2 and a severe error rate of 17.5%. These findings demonstrate that conventional aggregate metrics are insufficient for ensuring longitudinal decision safety in clinical settings.
📝 Abstract
Intraoperative anesthesia requires systems to interpret evolving multimodal evidence, recommend timely management, and revise decisions as patient states change, yet existing benchmarks usually isolate perception or single-point reasoning. We introduce AnesTRACE, an evaluation suite comprising AnesTRACE-Bench and AnesTRACE-Eval. Built from public perioperative datasets with anesthesiologist annotation, AnesTRACE-Bench evaluates Intraoperative Perception, Single-point Anesthesia Decision-Making, and Multi-step Anesthesia Decision-Making. AnesTRACE-Eval assesses open-ended responses through anesthesiologist-defined criteria for Clinical Correctness, Evidence Grounding, Task Completeness, and Safety, with Temporal Consistency for multi-step decisions; its domain-specific evaluator is trained by supervised fine-tuning and preference alignment on expert-reviewed judgments. Across more than 30 models, fine-grained visual grounding and intervention selection remain difficult: the leading model reaches only 32.2 mIoU for TEE visual grounding and retains a 17.5\% Major/Critical Safety Error Rate in multi-step management. Evaluator alignment with anesthesiologists improves across both training stages, while the best decision quality is accompanied by a 74.3-second P95 Latency. These results show that aggregate performance alone does not establish safe, timely longitudinal decision-making. We release our code at https://github.com/zjuDBxAI/AnesTRACE.
Problem

Research questions and friction points this paper is trying to address.

Intraoperative Anesthesia
Multimodal Perception
Multi-step Decision-Making
Benchmarking
Clinical Safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Perception
Multi-step Decision-Making
Benchmark Evaluation
Preference Alignment
Clinical Safety
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Ziwei Huang
Ziwei Huang
Zhejiang University
Multimodal LLMsAIGC
Qi Gao
Qi Gao
Department of Anesthesiology, The Second Affiliated Hospital of Zhejiang University School of Medicine
Z
Zhe Ji
School of Software, Central South University
Y
Yuanyuan Yao
Department of Anesthesiology, The Second Affiliated Hospital of Zhejiang University School of Medicine
F
Fengjiang Zhang
Department of Anesthesiology, The Second Affiliated Hospital of Zhejiang University School of Medicine
M
Min Yan
Department of Anesthesiology, The Second Affiliated Hospital of Zhejiang University School of Medicine
Zhongle Xie
Zhongle Xie
Zhejiang University
AI4DBML SystemsDB SystemsOLAP
G
Gang Chen
School of Software Technology, Zhejiang University