Beyond Component Testing: Validating Agentic AI Systems

📅 2026-07-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing verification methods for AI agent systems struggle to assess the reliability of multi-step decision trajectories in dynamic environments. Through a systematic literature review of 257 studies, this work constructs a five-dimensional verification taxonomy encompassing behavioral, safety, temporal, regulatory, and multi-agent aspects. Analysis of case studies from healthcare, industrial automation, and intelligent transportation reveals critical gaps in current research, particularly concerning temporal validity, runtime evidence maintenance, regulatory interpretability, and assurance in open multi-agent settings. The study proposes a lifecycle-oriented verification agenda and outlines four key directions: bounded autonomy specifications, adversarial trajectory generation, runtime monitoring, and auditable evidence structures—collectively offering a pathway toward context-aware, trajectory-level trustworthy verification.
📝 Abstract
Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation. This behavior stretches validation practice beyond component testing and one-shot input--output evaluation, because acceptable system behavior now depends on how decisions unfold over time and under changing environmental conditions. This survey synthesizes 257 papers spanning agent evaluation, software assurance, cyber-physical systems, runtime monitoring, and regulatory guidance in order to characterize the validation problem for agentic systems. The review is organized around a five-dimension taxonomy covering behavioral, safety, temporal, regulatory, and multi-agent concerns, and uses that taxonomy to map current approaches and expose recurrent coverage gaps. The analysis shows that behavioral evaluation is comparatively mature, while temporal validity, runtime evidence maintenance, regulatory legibility, and open-ended multi-agent systems assurance remain under-developed. Three cross-domain case studies (medical care, industrial operations, smart-mobility systems) provide operational illustrations of how the five taxonomy dimensions recur in safety-critical settings, grounded in the failure patterns documented in the reviewed literature. The paper concludes with a lifecycle-oriented research agenda centered on bounded-autonomy specifications, adversarial trajectory generation, runtime monitoring, and audit-ready evidence structures. The central claim is that trustworthy deployment of agentic AI depends on validating trajectories in context rather than assessing isolated components alone.
Problem

Research questions and friction points this paper is trying to address.

agentic AI
validation
behavioral trajectories
runtime monitoring
multi-agent systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

agentic AI validation
trajectory-based evaluation
runtime monitoring
bounded autonomy
multi-agent assurance
🔎 Similar Papers
No similar papers found.
F
Fabio Orazio Mirto
Department of Biomedical, Dental, and Morphological and Functional Imaging Sciences, University of Messina, A.O.U. Policlinico "G.Martino" - Via Consolare Valeria, Messina, 98125, Italy; Department of Engineering, University of Messina, Contrada di Dio, Sant'Agata, Messina, 98158, Italy
L
Luca D'Agati
Department of Engineering, University of Messina, Contrada di Dio, Sant'Agata, Messina, 98158, Italy
G
Giuseppe Tricomi
Department of Engineering, University of Messina, Contrada di Dio, Sant'Agata, Messina, 98158, Italy
Stefano Silvestri
Stefano Silvestri
Researcher, National Research Council of Italy (CNR) - Institute for High Performance Computing
Parallel computingNatural Language ProcessingDeep LearningMedical Informatic
Francesco Longo
Francesco Longo
Associate Professor of Computer Engineering, Università degli Studi di Messina
Computing ContinuumInternet of ThingsDistributed LedgersSmart CityAdditive Manufacturing
Antonio Puliafito
Antonio Puliafito
Università di Messina
distributed systemsperformance evaluationcloudcyber physical systemssmart cities
Giovanni Merlino
Giovanni Merlino
Associate Professor, Department of Engineering, University of Messina, Italy
Cloud ComputingInternet of ThingsSmart SensorsMobile Crowd SensingSoftware-Defined