Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了基于大语言模型的自主AI系统的安全风险,并提出了一种包含六个阶段的信任代理开发生命周期框架来缓解这些问题。
📝 Abstract
Agentic AI systems built on large language models can plan over multiple steps, use external tools, retain information in memory, and coordinate with other agents. These capabilities make them more useful than static language models, but they also introduce new security and operational risks. Untrusted content from websites, emails, documents, and databases can enter the same context as system instructions; persistent memory can carry compromised information across sessions; and access to external tools can turn an incorrect model response into a consequential real-world action. This article reviews the trustworthiness of agentic AI across five interconnected dimensions: safety and robustness, alignment and human oversight, transparency and auditability, privacy and data governance, and regulatory compliance. It organizes key failure modes, including indirect prompt injection, backdoor triggers, goal misgeneralization, memory contamination, and cross-session data leakage, into a unified taxonomy. It also examines major mitigation approaches, such as instruction hierarchies, context isolation, spotlighting, process-based supervision, constrained tool use, and privacy-preserving memory, while distinguishing techniques supported by empirical evidence from those that remain largely conceptual. Building on this analysis, we introduce the Trustworthy Agent Development Lifecycle (TADL), a six-phase framework covering specification, design, training, evaluation, deployment, and monitoring. For each phase, TADL identifies relevant trust activities, expected evidence, and risk-based decision gates. Although TADL has not yet been empirically validated, it provides a structured foundation for developing and evaluating more secure and accountable agentic systems. The article concludes by identifying gaps in current benchmarks and outlining priorities for future research.
Problem

Research questions and friction points this paper is trying to address.

Agentic AI
Security Risks
Operational Risks
Trustworthiness
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Trustworthy Agent Development Lifecycle
TADL
agentic AI systems
failure modes taxonomy
mitigation strategies
F
Fayeq Jeelani Syed
Department of Human-Centered Computing, Indiana University Indianapolis, 755 W. Michigan Street, Indianapolis, Indiana, 46202, USA
R
Rehan Ahmad
Department of Electrical and Computer Engineering, Purdue University Northwest, 2200 169th Street, Hammond, Indiana, 46323, USA
A
Ali Al Bataineh
Computer Information Science, Higher Colleges of Technology, Abu Dhabi, United Arab Emirates
A
Aakriti Adhikari
Department of Biomedical Informatics and Data Science, Yale School of Medicine, Yale University, 333 Cedar Street, New Haven, Connecticut, 06510, USA