🤖 AI Summary
Existing benchmarks struggle to evaluate the automation of healthcare workflows characterized by high policy density, multi-role coordination, and long-horizon interactions. This work proposes the first evaluation framework specifically designed for long-horizon automation of policy-intensive, role-complex, and irreversible enterprise processes. It introduces a high-fidelity simulation environment encompassing three core scenarios: prior authorization, utilization management, and care management. The environment integrates 87 MCP tools and over 1,290 documentation-based procedural manuals, requiring AI agents to complete end-to-end tasks through multi-role switching, multi-turn dialogue, and artifact generation. Among 30 agent configurations tested, the best-performing model achieved only a 28.0% task completion rate; under stricter evaluation criteria, none exceeded 20%, with single-session success rates as low as 3.8%, revealing significant limitations of current approaches in handling complex healthcare workflows.
📝 Abstract
End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions must be grounded in a large library of medical, insurance, and operational rules; Multi-role composition: a single task requires the agent to play multiple roles with handoffs; and multilateral interaction: intermediate workflow steps are multi-turn dialogs, such as peer-to-peer review and patient outreach. We introduce $χ$-Bench, a benchmark of long-horizon healthcare workflows across three domains: provider prior authorization, payer utilization management, and care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed via 87 MCP tools, which it must drive to a terminal status through tool calls and writing the role's artifacts, guided by a 1,290+ document managed-care operations handbook skill. Across 30 agent harness/models configurations, the best agent resolves only 28.0% of tasks, no agent clears 20% on strict pass^3, and executing all tasks in a single session slumps the performance to 3.8%. These results raise the hypothesis that similar gaps are likely to surface in other policy-dense, role-composed, irreversible enterprise domains.