Score
Designs and builds models and monitoring systems that detect or predict failure of an ongoing interaction or session (for example a dialogue) from partial inputs, producing early warning signals and estimates of whether or when the final outcome will be a failure. Analyses include selecting features and decision thresholds for timely alerts, training classifiers or regressors for session-level failure prediction, and evaluating performance across varying degrees of partial visibility and time horizons.
This work addresses the limitation of existing task-oriented dialogue system evaluation methods, which typically intervene only after clear failures have occurred, thereby hindering early mitigation. To enable timely intervention, the authors propose a multi-granularity early failure prediction approach that jointly models dialogue state utterances and belief state transition trajectories for the first time. Leveraging a dual-stream architecture to fuse these complementary signals, the method effectively identifies failure risk during the 25%–75% progress window of a dialogue—well before completion. The approach demonstrates consistent effectiveness under both real and generated belief states, significantly outperforming heuristic, classical, and single-stream baselines across multiple dialogue progress points. By providing actionable early warnings, it opens a critical window for system self-repair and proactive recovery.
This work addresses the challenge of early warning for failing interactions when only trajectory-level success/failure labels are available, where failure evidence is sparse and typically emerges late in the interaction. The authors propose a two-stage approach: first, a weakly supervised learning framework leveraging attention mechanisms to infer sparse turn-level failure signals from trajectory labels and estimate failure risk conditioned on partial interaction history; second, an α-STOP preference-conditioned stopping policy that enables dynamic, on-demand warnings. By explicitly modeling the sparse structure of failure evidence—rejecting the common assumption of uniformly distributed labels—the method achieves significant gains: failure indicators constitute merely 4.7–11.3% of turns and predominantly occur in the latter half of trajectories. Experiments show the risk predictor improves Pareto front performance by 1–10% over baselines, the full system yields 3–42% gains, and training costs are reduced by one to three orders of magnitude.
This study addresses the challenge of “silent failures” in long-running LLM agent systems, where errors are often masked as fluent and plausible yet incorrect narratives, impeding timely intervention. Through a longitudinal analysis of a personal assistant LLM agent continuously operating since March 2026 over an eight-week period, the authors conduct root-cause investigations of 22 incidents to propose the first taxonomy of five failure mechanisms specific to LLM agents and formally define the phenomenon of “fail-plausible” behavior. Leveraging a production-grade architecture—comprising 40 scheduled tasks, 8 LLM providers, tool governance agents, and a memory layer—alongside 4,286 unit tests, 827 governance checks, and manual retrospective audits, the study reveals that 70% of silent failures were only detectable by users, while retrospective auditing prevented 87% of recurrence but offered no preemptive mitigation. Failures exhibited latency up to 60 days and predominantly originated from inter-component gaps.
Large language models (LLMs) can still generate unsafe outputs in deployment, necessitating efficient real-time monitoring. This work proposes a lightweight online safety monitoring mechanism that integrates signals from an external verification model, threshold-based decision rules, and risk control theory to produce reliable alerts through calibrated thresholds. The approach features a simple architecture that avoids computationally intensive procedures yet achieves detection performance on par with state-of-the-art sequential hypothesis testing methods across mathematical reasoning and red-teaming benchmarks. By combining practical efficiency with theoretical guarantees, the proposed method offers a viable solution for real-world LLM safety monitoring.
Naval systems frequently exhibit anomalous behaviors due to wear, misuse, or component failures—challenges that hinder timely detection and precise remediation. To address this, we propose a predictive-diagnostic closed-loop framework that tightly integrates the existing failure prediction system PREVENT with a newly designed responsive troubleshooting module, REACT. Methodologically, the framework synergizes multi-source time-series anomaly detection with domain-knowledge-driven fault-isolation process modeling, enabling end-to-end automation—from anomaly alerting and root-cause localization to actionable remediation recommendations. Evaluated on operational shipboard systems deployed by Fincantieri, the framework reduces mean time to fault localization by 42%, significantly improves operational response efficiency, and demonstrates strong generalizability across diverse industrial domains.
This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).
This work addresses the challenge of unreliable sensor readings in industrial inspection robots caused by occlusions, limited viewpoints, or environmental anomalies, which hinder real-time task status assessment. The authors propose a hybrid framework that integrates supervised fault classification with unsupervised anomaly detection, uniquely combining conformal prediction and world models to enable policy-agnostic, distribution-free early discrimination among three states—success, known faults, and out-of-distribution anomalies—using compressed video inputs. The approach facilitates training data quality evaluation and model feedback, achieving over 90% recognition accuracy on both office and industrial instrument inspection datasets. It outperforms human observers in decision speed and has been successfully deployed on a Boston Dynamics Spot robot for real-time operation.
This work addresses the limited trust in existing online fault prediction methods due to their poor interpretability and susceptibility to workload-specific noise. To overcome these challenges, the authors propose a multi-view complementary and interpretable framework tailored for Linux systems, integrating consensus-based feature selection, temporal onset analysis, subsystem-level causal inference, and multiple diagnostic mechanisms. The framework enables cross-workload evaluation under frozen model conditions. Experimental results demonstrate that the system achieves a detection rate of 91–94% on unseen workloads with a false positive rate below 1%. Furthermore, the study reveals key insights: detection is more robust than diagnosis, warning lead time is highly dependent on fault patterns, and unseen fault patterns cannot be accurately diagnosed solely based on related ones—highlighting the strong sensitivity of diagnosis to specific workloads.
This work addresses the challenge of predicting remaining useful life in complex manufacturing environments where existing models—reliant on fixed, known failure modes and labeled data—struggle to handle unknown, unlabeled faults arising in high-mix or adaptive production settings. To overcome this limitation, the authors propose a Bayesian nonparametric framework that integrates a Dirichlet process mixture model with a neural network–based prediction module, enabling joint online learning of fault mode discovery and remaining useful life estimation through an iterative feedback mechanism. The approach dynamically expands, merges, or infers novel fault modes without requiring static assumptions about their number or form. Evaluated on both synthetic and real-world aircraft engine datasets, the model demonstrates robustness and competitive or superior performance compared to state-of-the-art methods, highlighting its suitability for digital twin–enabled health management in dynamic manufacturing systems.