CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

📅 Unknown Date
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing benchmarks struggle to evaluate the automation of healthcare workflows characterized by high policy density, multi-role coordination, and long-horizon interactions. This work proposes the first evaluation framework specifically designed for long-horizon automation of policy-intensive, role-complex, and irreversible enterprise processes. It introduces a high-fidelity simulation environment encompassing three core scenarios: prior authorization, utilization management, and care management. The environment integrates 87 MCP tools and over 1,290 documentation-based procedural manuals, requiring AI agents to complete end-to-end tasks through multi-role switching, multi-turn dialogue, and artifact generation. Among 30 agent configurations tested, the best-performing model achieved only a 28.0% task completion rate; under stricter evaluation criteria, none exceeded 20%, with single-session success rates as low as 3.8%, revealing significant limitations of current approaches in handling complex healthcare workflows.
📝 Abstract
End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions must be grounded in a large library of medical, insurance, and operational rules; Multi-role composition: a single task requires the agent to play multiple roles with handoffs; and multilateral interaction: intermediate workflow steps are multi-turn dialogs, such as peer-to-peer review and patient outreach. We introduce $χ$-Bench, a benchmark of long-horizon healthcare workflows across three domains: provider prior authorization, payer utilization management, and care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed via 87 MCP tools, which it must drive to a terminal status through tool calls and writing the role's artifacts, guided by a 1,290+ document managed-care operations handbook skill. Across 30 agent harness/models configurations, the best agent resolves only 28.0% of tasks, no agent clears 20% on strict pass^3, and executing all tasks in a single session slumps the performance to 3.8%. These results raise the hypothesis that similar gaps are likely to surface in other policy-dense, role-composed, irreversible enterprise domains.
Problem

Research questions and friction points this paper is trying to address.

healthcare workflows
policy density
multi-role composition
multilateral interaction
long-horizon automation
Innovation

Methods, ideas, or system contributions that make the work stand out.

policy density
multi-role composition
multilateral interaction
long-horizon workflows
healthcare automation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Haolin Chen
actA V A.ai
D
Deon Metelski
actA V A.ai
L
Leon Qi
actA V A.ai
T
Tao Xia
actA V A.ai
J
Joonyul Lee
actA V A.ai
S
Steve Brown
actA V A.ai
K
Kevin Riley
actA V A.ai
F
Frank Wang
actA V A.ai
T
T. Y. Alvin Liu
Johns Hopkins Medicine
H
Hank Capps
Wellstar Health System
Zeyu Tang
Zeyu Tang
Postdoctoral Scholar, Stanford University
Trustworthy AICausalityComputational Justice
Xiangchen Song
Xiangchen Song
Carnegie Mellon University
Machine LearningCausalityData Mining
Lingjing Kong
Lingjing Kong
Carnegie Mellon University
Machine Learning
Fan Feng
Fan Feng
University of California San Diego & MBZUAI
Machine LearningReinforcement LearningRepresentation Learning
Tianyi Zeng
Tianyi Zeng
Shanghai Jiao Tong University; Purdue University
Intelligent vehicleRoboticsMachine learning
Zhiwei Liu
Zhiwei Liu
Research Scientist, Salesforce
AI AgentMulti-Agent SystemRecommender SystemGraph Mining
Zixian Ma
Zixian Ma
University of Washington
Multi-modal models and agentshuman-agent interaction and collaboration
Hang Jiang
Hang Jiang
MIT
Large Language ModelsNatural Language ProcessingHuman-AI Interaction
F
Fangli Geng
Brown University
Yuan Yuan
Yuan Yuan
Assistant Professor in Computer Science at Boston College, previously at MIT.
Machine LearningComputer VisionMedical AIArtificial Intelligence
Chenyu You
Chenyu You
Assistant Professor, Stony Brook University
Machine LearningAI for HealthComputer VisionMedical Image AnalysisMultimedia
Q
Qingsong Wen
University of Oxford
Hua Wei
Hua Wei
School of Computing and Augmented Intelligence, Arizona State University
Data MiningMachine LearningReinforcement Learning
Yanjie Fu
Yanjie Fu
Associate Professor at School of Computing and AI, Arizona State University
Artificial IntelligenceAI4DataSpatiotemporal IntelligenceSim2DecisionMultimodal Reasoning
Yue Zhao
Yue Zhao
Assistant Professor of Computer Science, University of Southern California
Anomaly DetectionOut-of-Distribution DetectionTrustworthy AIAI for ScienceML Systems