๐ค AI Summary
This study addresses the ambiguity surrounding the definition and scope of AI agents in healthcare and the lack of effective evaluation frameworks for clinical translation. Through a systematic evidence mapping of 557 studies, it establishes a clear conceptual boundary for medical AI agents and synthesizes their architectural designs and implementation strategies across key clinical tasksโincluding medical question answering, medical image interpretation, and electronic health record analysis. The work proposes a novel clinical evaluation framework emphasizing auditability, interoperability, and prospective validation. It reveals that current research predominantly relies on retrospective data and public benchmarks, while largely neglecting systematic assessments of safety, reliability, and impact on clinical workflows, thereby offering critical guidance for the future development and translational deployment of healthcare AI agents.
๐ Abstract
Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and undertake multistep clinical tasks that require planning, tool use, memory, iterative correction, and coordination among specialized agents. However, the scope of agentic AI in medicine remains unsettled, and current evaluation practices are not yet aligned with the requirements of clinical use. We conducted a scoping review with systematic evidence mapping across five electronic sources, screened 1,649 exportable records, and provisionally included 557 unique studies that met predefined criteria for goal-directed task execution, tool use, interaction with external resources, feedback-based refinement, or multi-agent collaboration. The included studies describe single agents that use external tools, workflows supported by retrieval and external knowledge, multimodal agents, and multi-agent systems applied to medical question answering, image interpretation, electronic health record analysis, drug safety, and clinical trial prediction. The evidence base remains dominated by public benchmarks, simulated settings, retrospective datasets, and small-scale expert evaluation. Process reliability, evidence traceability, uncertainty, safety, workflow impact, and external validity are evaluated less consistently. Clinical translation will depend on clearer definitions, reproducible evaluation, auditable oversight, interoperable system design, and prospective validation in real-world clinical workflows.