Reliability Engineering for AI Systems: Challenges, Methods, and Directions

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that traditional benchmarks struggle to quantify the broad reliability of increasingly autonomous AI systems across retrieval, memory, and reasoning stages. To this end, it adapts reliability engineering methodologies, including Failure Mode and Effects Analysis (FMEA) and accelerated testing, to the AI domain, constructing a four-tier fault diagnosis framework spanning from components to governance. By integrating the NIST risk management framework with SMART statistical methods and Test, Evaluation, Verification, and Validation (TEVV) processes, the work establishes a full-lifecycle evidence chain. Validated through CNN adversarial testing and autonomous driving case studies, this research facilitates a paradigm shift from evaluating isolated output correctness to ensuring system-level operational reliability. Ultimately, it lays a practical foundation for AI reliability engineering while revealing the necessity of novel metrics tailored for self-evolving systems.
📝 Abstract
AI reliability concerns whether an AI system performs its intended function dependably over a stated period and under stated operating conditions, with stated evidence. As these systems become more autonomous, that function includes more than a correct output. Retrieval, memory, tool use, permissions, human oversight, and interactions among systems must operate consistently and safely, and, for generative systems, so must the reasoning process that produces the output. Average benchmark accuracy measures capability; it does not quantify this broader reliability claim. This paper adapts established reliability engineering methods, from failure definitions and operational envelopes to FMEA, accelerated testing, field monitoring, and reliability growth, to AI systems. A four-level diagnostic framework classifies failures as component, operational-loop, agentic-conduct, or network and governance failures. Test, evaluation, verification, and validation (TEVV), sequential monitoring, and FRACAS create and refresh evidence. SMART provides statistical guidance for measurement, analysis, assessment, and test planning; the NIST AI Risk Management Framework provides organizational guidance for governance, evaluation, monitoring, and mitigation. Three cases illustrate the program: adversarial testing of a convolutional neural network, perception-error propagation, and autonomous-vehicle disengagements. Established reliability engineering provides a usable foundation; new measurements and safety guardrails are still needed as these systems are self-evolving.
Problem

Research questions and friction points this paper is trying to address.

AI reliability
reliability engineering
autonomous systems
failure classification
safety guardrails
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reliability Engineering
Four-level Diagnostic Framework
FMEA
TEVV
SMART Framework
🔎 Similar Papers
No similar papers found.