MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current evaluations of medical large language models are often confined to exam-based knowledge or isolated tasks, failing to capture the longitudinal, multimodal, and safety-critical nature of real-world cardiovascular clinical workflows. To address this gap, this work introduces MyoCardBench, a real-world benchmark spanning the entire care continuum, comprising 2,263 tasks derived from de-identified electronic health records and examination data. The benchmark was annotated by 16 cardiologists and validated by two senior experts, enabling evaluation of 15,841 zero-shot outputs from seven large language models. MyoCardBench represents the most comprehensive assessment to date of cardiovascular clinical scenarios, incorporating a dual-dimensional evaluation framework that measures both key point coverage and overall clinical quality, including challenging tasks such as ECG interpretation and ethical reasoning. Results show that GPT-5.4 achieves the best overall performance (macro-average 62.55) yet still exhibits notable deficiencies in complex clinical reasoning.
📝 Abstract
Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, multimodal, and safety-critical workflow of cardiovascular care. Objective: To develop MyoCardBench, a real-world benchmark spanning the cardiovascular care continuum, and assess LLM performance across clinical dimensions and specialist tasks. Methods: MyoCardBench includes 2,263 items from 13 task-specific datasets derived from de-identified cardiovascular records and examination data. Sixteen cardiology physicians conducted annotation and reference construction, followed by cross-review from two senior cardiologists. Seven LLMs generated 15,841 outputs under standardized zero-shot settings. Open-ended tasks were evaluated using key-point coverage and holistic clinical quality, while CardioEthics was scored by accuracy. Results: GPT-5.4 achieved the highest macro-average (62.55) and item-weighted mean (62.19), followed by Gemini 3.1 Pro (59.95) and Qwen 3.6 27B (59.72). GPT-5.4 ranked first in all three dimensions. CardioAuxReport performed best (86.38), whereas CardioECGRead (17.25) and CardioEthics (17.34) were lowest. The largest gaps between holistic clinical quality and key-point coverage occurred in CardioComm (52.71), CardioEmergRescue (52.05), and CardioTreatPlan (48.80). Conclusions: To our knowledge, MyoCardBench is the largest real-world, multi-task benchmark for LLM evaluation across the cardiovascular care continuum and offers the broadest coverage of clinically authentic cardiology scenarios reported to date. It provides a rigorous framework for identifying model strengths, clinically important omissions, and priorities for future development.
Problem

Research questions and friction points this paper is trying to address.

large language models
cardiovascular care
clinical benchmark
real-world data
medical AI evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

real-world benchmark
cardiovascular care continuum
clinically authentic evaluation
multimodal clinical tasks
CardioEthics
🔎 Similar Papers
No similar papers found.
Xiao Li
Xiao Li
Shanghai Jiaotong University
roboticsreinforcement learningformal methods
M
Mouxiao Bian
Shanghai Artificial Intelligence Laboratory, Shanghai, China
Z
Zhaodi Wu
Minhang Hospital, Fudan University, Shanghai, China
S
Sijie Ren
Shanghai Artificial Intelligence Laboratory, Shanghai, China
J
Juechen Chen
Department of Cardiology, Zhongshan Hospital Fudan University, Shanghai, China; Shanghai Institute of Cardiovascular Diseases, National Clinical, Shanghai, China; Research Center for Interventional Medicine, Shanghai, China
Lu Lu
Lu Lu
University of Chinese Academy of Sciences
Wireless CommunicationsPhysical-layer Network CodingSoftware-defined Radio
J
Jingru Ding
Shanghai Artificial Intelligence Laboratory, Shanghai, China
Yun Zhong
Yun Zhong
Imperial College London
Human Motion AnalysisMachine LearningMultimediaSelf-Supervised LearningPrompt Tuning
J
Jie Xu
Shanghai Artificial Intelligence Laboratory, Shanghai, China
Y
Yixiu Liang
Department of Cardiology, Zhongshan Hospital Fudan University, Shanghai, China; Shanghai Institute of Cardiovascular Diseases, National Clinical, Shanghai, China; Research Center for Interventional Medicine, Shanghai, China
J
Junbo Ge
Department of Cardiology, Zhongshan Hospital Fudan University, Shanghai, China; Shanghai Institute of Cardiovascular Diseases, National Clinical, Shanghai, China; Research Center for Interventional Medicine, Shanghai, China