PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current evaluations of medical AI systems are largely confined to static question-answering or clinician-facing tasks, failing to adequately assess the safety and efficacy of patient-facing health agents in real-world clinical workflows. This work proposes the first dynamic evaluation framework tailored to patient-side interactions, which simulates patient dialogues, integrates a sandboxed medical tool environment, and employs an LLM-based jury mechanism to automatically score performance across six dimensions against over 100 clinical criteria. The framework implements a dialogue-agnostic, multidimensional assessment protocol validated by licensed physicians and evaluates ten leading models across 1,200 clinical scenarios. Results reveal that despite achieving an average score of 4.25 out of 5, state-of-the-art models exhibit critical shortcomings in key areas such as triage accuracy and clinical safety, underscoring the necessity of dynamic evaluation for uncovering workflow-related risks and latent safety hazards.
📝 Abstract
Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in this domain warrant evaluation against the same risks. Current benchmarks focus on medical knowledge, assessed through isolated question-answering or clinician-facing tasks. PatientAgentBench benchmarks patient-facing agentic healthcare; it evaluates a foundation model, wrapped in an agent with a sandbox of healthcare tools, conversing with a simulated patient. Each conversation is scored by an LLM-as-a-Jury across six dimensions via over a hundred conversation-agnostic, clinician-grounded criteria. To validate alignment, licensed clinicians annotated shared conversations, yielding 79-93% adjacent agreement between jury and expert raters, on par with or exceeding clinician inter-rater agreement. We benchmarked 10 models across four families on the same 1,200 scenarios and found clinical gaps. Triage quality is the most discriminating dimension: pass rates rise from 32% for the weakest models to 88% for the strongest, with agents often acting on administrative requests without clinical screening. Clinical safety and workflow accuracy follow the same pattern: the weakest models fail often, fabricating unexecuted actions, while frontier models fail on only 1-3% of cases, from unverified tool outputs and omitted crisis resources in an emergency. More capable models narrow these gaps but do not close them; the strongest scores only 4.25 of 5 overall. These failures surface only in sustained, tool-using conversations against realistic patient records, confirming that static benchmarks are insufficient as healthcare agentic systems gain autonomy. We release the framework as a reproducible, clinician-validated evaluation standard to help the field close this gap.
Problem

Research questions and friction points this paper is trying to address.

patient-facing AI
healthcare agents
clinical safety
benchmarking
diagnostic errors
Innovation

Methods, ideas, or system contributions that make the work stand out.

PatientAgentBench
healthcare AI agents
LLM-as-a-Jury
clinical safety evaluation
agentic benchmarking
Korosh Vatanparvar
Korosh Vatanparvar
Senior Applied Scientist, Amazon
HealthDeep LearningGenerative AI
Ashutosh Joshi
Ashutosh Joshi
Amazon.com
M
Maria Xenochristou
Amazon Health AI
M
Mohammad Abuzar Hashemi
Amazon Health AI
P
Prasad Kasu
Amazon Health AI
D
Deepak Bansal
Amazon Health AI
Daniel Lopez-Martinez
Daniel Lopez-Martinez
Harvard University & Massachusetts Institute of Technology
Medical Engineering and Medical Physics
A
Anchal Nema
Amazon Health AI
R
Ramya Ganesan
Amazon Health AI
W
Will Kimbrough
Amazon Health AI
A
Alex Woody
Amazon Health AI
Y
Yadunandana Rao
Amazon Health AI
D
Dilek Hakkani-Tur
Amazon Health AI
W
Wilko Schulz-Mahlendorf
Amazon Health AI