🤖 AI Summary
This study addresses the challenge of “silent failures” in long-running LLM agent systems, where errors are often masked as fluent and plausible yet incorrect narratives, impeding timely intervention. Through a longitudinal analysis of a personal assistant LLM agent continuously operating since March 2026 over an eight-week period, the authors conduct root-cause investigations of 22 incidents to propose the first taxonomy of five failure mechanisms specific to LLM agents and formally define the phenomenon of “fail-plausible” behavior. Leveraging a production-grade architecture—comprising 40 scheduled tasks, 8 LLM providers, tool governance agents, and a memory layer—alongside 4,286 unit tests, 827 governance checks, and manual retrospective audits, the study reveals that 70% of silent failures were only detectable by users, while retrospective auditing prevented 87% of recurrence but offered no preemptive mitigation. Failures exhibited latency up to 60 days and predominantly originated from inter-component gaps.
📝 Abstract
LLM agent systems increasingly run as long-lived autonomous runtimes: scheduling jobs, calling tools, maintaining memory, and pushing results to humans. We present a longitudinal study of silent failures in one such system: a personal-assistant agent runtime in continuous production since March 2026, with roughly 40 scheduled jobs, 8 LLM providers, a tool-governance proxy, and a knowledge-base memory plane, defended by 4,286 unit tests and 827 governance checks. Over eight weeks we documented 22 incidents with full root-cause postmortems, in which one meta-pattern -- a failure whose error signal never reaches a human in actionable form -- manifested at least 28 times. We derive a five-class, mechanism-oriented taxonomy: (A) environment and platform quirks, (B) design-assumption mismatches, (C) error swallowing and dilution, (D) chained hallucination and fabrication, (E) operational omission and forensic blind spots. Class D is unique to LLM systems and the most dangerous: the system does not merely fail to report an error -- the LLM transforms it into fluent, plausible narrative delivered to the user. We term this fail-plausible: gray failure's differential observability escalated -- the observer is not just blind, it is convincingly lied to by the failure itself. Three findings: about 70% of silent failures were caught by human user-view observation, not tests or audits; a retrospective audit of 15 incidents found 0% ex-ante prevention but 87% regression blocking -- audits are regression engines, not prediction engines; incident latency (13 hours to 60 days) tracks failure mechanism, not code complexity -- the longest-lived failures lived in the seams between components, where no test runs. We describe the resulting defense framework and distill design principles for agent systems whose failures are loud, attributable, and boring. All postmortems and artifacts are public.