EHR-RobustGym: Benchmarking and Training Agents for Robust Clinical Reasoning

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the tendency of clinical agents to overlook missing evidence and generate ungrounded responses when operating on noisy electronic health records (EHRs). We construct an interactive environment based on MIMIC-IV and introduce the first scalable interaction benchmark supporting multi-granularity noise testing. By generating clean–noisy data pairs through SQL/Python execution with result verification, we establish a closed-loop optimization pipeline from evaluation to training via supervised fine-tuning and reinforcement learning, yielding a robust clinical reasoning framework. Our experiments reveal significant robustness deficiencies in existing models, while the proposed training approach substantially improves performance and demonstrates strong generalization across five external EHR benchmarks.
📝 Abstract
In hospital workflows, electronic health records (EHRs) are often noisy, and may not contain the evidence needed to confirm events or measurements referenced in a clinical query. Even when database retrieval succeeds, clinical agents can overlook such discrepancies and return plausible but unsupported answers. We introduce EHR-RobustGym, a scalable and interactive environment for evaluating and training robust clinical agents grounded in noisy EHRs. Built on MIMIC-IV hospital records (365K patients, 31 tables, and over 500M records), EHR-RobustGym comprises 5,486 Clean-Noise pairs spanning six clinical intents and both patient-level and population-level queries. The pairs test robustness to Record-level, Value-level, and Query-level noise, while interactive SQL/Python execution and outcome verification support trajectory collection and training. Evaluating multiple LLMs reveals substantial robustness gaps: average task success across proprietary and large-scale open-weight models drops from 62.2% on Clean questions to 37.9% on Noise questions. At k=4, pass^k consistency falls below 50% for most evaluated models, exposing instability in clinical task completion. Supervised fine-tuning and reinforcement learning in EHR-RobustGym improve performance, with gains generalizing to five external EHR benchmarks. Together, these results position EHR-RobustGym as a testbed for evaluating and improving the evidence-grounded robustness of clinical agents.
Problem

Research questions and friction points this paper is trying to address.

Electronic Health Records
Clinical Reasoning
Robustness
Noisy Data
Clinical Agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

EHR-RobustGym
Clinical Reasoning
Robustness Benchmarking
Interactive Environment
Reinforcement Learning
🔎 Similar Papers
No similar papers found.
Y
Yitong Qiao
Zhejiang University; Ant Healthcare, Ant Group
Y
Yancheng Jin
Zhejiang University; Ant Healthcare, Ant Group
L
Lei Liu
Zhejiang University; Ant Healthcare, Ant Group
Yue Shen
Yue Shen
Ant Group
recommend user growth
J
Jian Wang
Ant Healthcare, Ant Group
Jinjie Gu
Jinjie Gu
ant group
机器学习,推荐
Zhixuan Chu
Zhixuan Chu
Associate Professor, Zhejiang University; Alibaba Group; Ant Group