Score
Designs and evaluates pipelines, distributed architectures, and agent orchestration protocols that maintain correct, available, and safe operation in the presence of component failures or malfunctions. Builds redundant and isolated processing stages, cross-checking and adversarial-review mechanisms, error-detection and recovery protocols, and explicit, bounded human‑intervention paths to enable graceful degradation and controlled recovery.
This study addresses reliability challenges in composite AI systems arising from cascading errors, silent degradation, and coordination failures at component boundaries. Drawing on 150 production incidents, it presents the first systematic fault taxonomy and proposes a catalog of resilience patterns encompassing five failure categories, including retrieval and generation faults. The effectiveness of mitigation strategies—specifically circuit breakers, output quality gates, and component isolation—is validated through fault injection experiments. Results demonstrate that circuit breakers reduce cascading propagation by 89%, quality gates detect 73% of silent degradations, and combining multiple patterns decreases mean time to recovery (MTTR) by 71%. These findings provide practitioners with an empirically grounded resource for reliability engineering in complex AI architectures.
To address the frequent runtime anomalies in large language model (LLM)-based agent systems and the absence of systematic operational methodologies, this paper presents the first comprehensive survey on AgentOps—the operations of intelligent agent systems. Through systematic literature review and anomaly taxonomy modeling, we formally define internal anomalies (e.g., reasoning hallucinations, tool invocation failures) and external anomalies (e.g., API outages, environmental changes). Building upon this taxonomy, we propose a full-lifecycle AgentOps framework encompassing monitoring, anomaly detection, root-cause analysis, and autonomous recovery. This work establishes the first structured conceptual foundation for agent operations, fills a critical theoretical gap in the field, identifies key open challenges, and provides an extensible methodological basis and evolutionary roadmap for both academic research and industrial deployment.
This work addresses the limitation of existing benchmarks, which focus solely on accuracy in multi-agent orchestration tasks while neglecting fine-grained diagnosis of failure origins and recovery capabilities. The authors propose a reproducible fault-injection framework to systematically evaluate failure modes, task decomposition quality, and recovery mechanisms within templated enterprise workflows. They introduce two novel metrics: “cascade radius” and failure-mode-specific recovery rates, and employ controlled probes to analyze recovery behavior across different fault types. Experimental results demonstrate that intent-based reasoning routing achieves 100% recovery under adversarial conditions, significantly outperforming keyword-based routing; tool-related failures are fully recoverable, whereas semantic failures prove largely irrecoverable; and cascade radius increases with workflow depth.
This work addresses the reliability of AI systems under unexpected failures and adversarial attacks. We systematically adapt Byzantine Fault Tolerance (BFT)—a foundational paradigm from distributed systems—to AI safety, introducing the novel conceptual analogy that “malicious AI modules correspond to Byzantine nodes.” Based on this, we formalize a component-level AI failure model and design a multi-agent consensus verification framework integrating distributed consensus protocols, behavioral consistency checking, redundant heterogeneous model arbitration, and controlled fault-injection testing. Evaluated across multiple high-stakes AI decision-making tasks, our architecture achieves a 99.2% anomaly detection rate, substantially enhancing robustness and trustworthiness against both adversarial perturbations and internal component failures. The proposed approach establishes a verifiable, scalable, and principled new paradigm for AI safety.
This work addresses the trade-off between reliability and cost in large language model (LLM) agents for tool invocation: fully LLM-driven routing incurs high computational overhead, while static workflows are prone to silent failures under compound tool errors. To overcome this, we propose the Self-Healing Router architecture, which introduces runtime fault tolerance into graph-based tool-calling systems for the first time. By modeling control flow as a weighted graph routing problem, our approach employs parallel health monitoring to dynamically assess tool status and leverages Dijkstra’s algorithm for deterministic shortest-path rerouting, invoking the LLM only when necessary. Evaluated across 19 scenarios, the method achieves accuracy comparable to ReAct while reducing control-plane LLM calls by 93% (from 123 to 9) and completely eliminating silent failures, thereby enabling binary observability and efficient autonomous recovery.
研究通过构建每个动作的真实值来解决NetOps代理在数据中心网络修复中的长期可靠性问题,利用这些信号提高代理预测行动风险和进展的能力。
This work addresses the frequent disruptions in data and AI pipelines caused by data anomalies, schema changes, or infrastructure failures, which existing self-healing solutions often mitigate at high cost, with vendor lock-in, or with limited adaptability for small-to-medium teams. The paper proposes an open-source, vendor-agnostic autonomous remediation architecture that integrates monitoring, metadata, historical incident logs, a policy engine, AI-driven root cause analysis, and controlled repair mechanisms to automatically detect, diagnose, remediate, and validate pipeline issues. Its key contribution lies in delivering a unified, portable end-to-end reference architecture that systematically consolidates fragmented capabilities without reliance on proprietary platforms, substantially reducing manual intervention and enhancing pipeline resilience across diverse contexts such as data engineering, MLOps, and software delivery.
This work addresses the security risks propagated across stages in multi-stage workflows involving large language model (LLM) agents, a challenge inadequately handled by existing approaches that focus on isolated stages without holistic coordination. To bridge this gap, the authors introduce the abstraction of Stage-Specific Safety Skills, which modularizes heterogeneous safety mechanisms into reusable and composable components. They further develop an automated transformation pipeline and a community-driven safety skill repository. Building upon this foundation, they propose the $S^3$ (Stage-Specific Safety Skills) multi-stage defense framework, wherein guardian agents orchestrate stage-specific skills to enable end-to-end risk detection and mitigation. Experimental results demonstrate that $S^3$ significantly outperforms current methods in both safety effectiveness and task utility, highlighting its potential for constructing trustworthy LLM agent systems.
This work addresses the limitations of traditional expert-manual-based cybersecurity response methods, which struggle to adapt to dynamic attack scenarios and evolving recovery objectives, as well as the instability of existing large-model approaches in long-horizon tasks. The authors propose an end-to-end agent planning framework that innovatively models event states using a graph structure (Graph-as-State), incorporates a phase-aware agent routing mechanism, and establishes a verifiable experience reuse loop to guide action selection and state updates. The system integrates multi-agent large language models with experience retrieval augmentation and execution feedback verification, enabling dynamic, stable, and evolvable response planning within a Docker-based network range simulation environment. Experimental results demonstrate that the proposed method achieves a normalized defense score of 0.94 across 100 simulated scenarios, representing a 9.5% improvement over the strongest baseline.
Existing reliability metrics struggle to effectively detect latent failures in agent networks caused by stale information, redundant operations, partial execution, or silent degradation. To address this limitation, this work proposes the Reliability Assurance Intelligence (RAI) architecture, which introduces, for the first time, a scope-oriented reliability mechanism that unifies the modeling of correctness, auditability, and recoverability of intent execution. The architecture captures assurance requirements through service reliability profiles, persists critical runtime states via context capsules, and integrates generic runtime functions with deterministic service lifecycle management. Experimental results demonstrate that RAI accurately captures and validates essential states in agent lifecycle management scenarios, significantly enhancing the system’s capability to assure reliability.
This work addresses the challenges of cumulative error and reliability faced by embodied agents performing long-horizon tasks under resource constraints and environmental uncertainty. The authors propose a distributed fault-tolerant architecture leveraging edge–cloud collaboration, redefining reliability as system-level fault tolerance. Their approach features a two-tiered fault-tolerance mechanism: within individual agents, fault-tolerant alignment mitigates error propagation, while across heterogeneous agents, a semi-formal language protocol enables robust coordination. In contrast to conventional single-round, zero-error optimization paradigms, this framework substantially enhances the robustness of long-term task execution and offers a scalable engineering pathway for reliable collaboration among heterogeneous embodied agents in industrial settings.
研究解决了自主代理执行任务时遇到的可靠性问题,通过分析生产平台中的故障案例,提出了基于代理委托而非消息传递的七个可靠性原语。