Score
Designs and applies quantitative metrics and analysis methods that estimate the portion of performance degradation or errors attributable to memory failures (forgetting) as distinct from poor decisions or execution. Builds statistical decompositions and residual-error measures that disentangle memory loss from other error sources and produce a memory-gap metric for comparing agents, interventions, or models.
This work addresses the challenges faced by foundational agents in long-horizon, dynamic, and user-dependent environments—particularly context explosion and sustained information management—which necessitate efficient memory mechanisms to enhance practical utility. The paper proposes the first unified three-dimensional framework that integrates internal and external memory, five cognitive mechanisms, and a dual-agent/user-centric perspective, offering a systematic structuring of research on agent memory. Drawing on a comprehensive review of hundreds of studies published before 2026, the framework synthesizes insights from memory modeling, cognitive science, and agent architecture to clarify memory operation strategies, evaluation benchmarks, and learning methodologies. It not only delineates structured pathways in existing research but also identifies key open problems, thereby providing a theoretical foundation and directional guidance for the future design of intelligent agent memory systems.
When deploying machine learning models across heterogeneous environments, performance degradation often exhibits subgroup-specific heterogeneity—yet existing methods either explain only mean-level distributional shifts or isolate vulnerable subgroups without jointly identifying *where* degradation occurs and *why* it arises. This paper introduces SHIFT, the first hierarchical inference framework that unifies subgroup scanning, hierarchical causal inference, variable subset sensitivity analysis, and interpretable shift attribution. SHIFT simultaneously enables precise identification of degraded subgroups and disentanglement of underlying causes—distinguishing covariate shift from outcome shift. Evaluated on real-world deployments, SHIFT generates human-interpretable attributions of performance degradation and guides targeted interventions: it significantly improves performance for affected subgroups while avoiding negative transfer to others.
This work addresses the limitation of traditional accuracy metrics in continual learning, which only provide a binary indication of catastrophic forgetting and fail to capture its internal structure. The authors propose six continuous distribution-based measures derived from softmax outputs—such as ground-truth label rank, prediction confidence, and distributional divergence—that enable fine-grained quantification of forgetting as a value within the [0,1] interval. These metrics reveal semantic differences even when accuracy drops to zero, without requiring modifications to the training procedure. Leveraging these measures, the study introduces a weighted experience replay mechanism and a trend-slope-based sample selection strategy. Experiments demonstrate that the proposed approach reduces forgetting by 1.3% on CIFAR-100 and by 7.7% on TinyImageNet compared to baseline methods, significantly outperforming existing techniques that rely solely on accuracy trends.
Traditional model evaluation relies on single-point metrics, failing to characterize performance stability and uncertainty. This paper proposes a small-sample (10–25 runs) uncertainty quantification framework tailored for high-reliability scenarios. It constructs empirical distributions of performance metrics via repeated stochastic experiments—encompassing random data splits, parameter initializations, and hyperparameter perturbations—and robustly estimates confidence intervals for metric quantiles using bias-corrected nonparametric bootstrap combined with quantile regression. To our knowledge, this is the first systematic approach enabling reliable confidence interval estimation for diverse metrics—including accuracy, F1-score, and MAE—in both classification and regression tasks under small-sample regimes. The method achieves high coverage (>90%) while maintaining narrow interval widths, thereby significantly improving robustness in model selection and enhancing decision-making credibility across multiple benchmark datasets.
This work addresses the misalignment between offline evaluation metrics and online performance objectives in industrial applications by establishing a unified theoretical framework that systematically quantifies the relationships among diverse evaluation metrics for the first time. By introducing the concepts of Bayes-optimal sets and regret transfer mechanisms, the study reveals structural asymmetries among metrics and provides a principled classification and relational modeling of metrics with varying mathematical forms. Theoretically characterizing metric consistency and transferability, this research offers novel insights and a methodological foundation for designing offline evaluation systems that are aligned with online objectives and backed by rigorous theoretical guarantees.
This work addresses core challenges in behavioral metric learning for deep reinforcement learning: the theory-practice gap, difficulty in evaluating metric quality, and unclear performance attribution. We propose an isometry-preserving, noise-robust policy-invariant representation learning framework. Evaluated systematically across 20 state-based and 14 pixel-based tasks under 370 noise configurations, our approach introduces the denoising factor—a novel metric quantifying encoder robustness—and designs an isolated metric estimation protocol to decouple learning effects. We further release the first open-source, modular benchmark library for behavioral metric learning. Experiments demonstrate that performance gains stem primarily from structural constraints in the latent space—not merely denoising—revealing systematic impacts of key design choices on generalization and stability. All methods achieve an average 23% improvement in policy robustness on high-interference tasks.
This study addresses the challenge of “silent failures” in long-running LLM agent systems, where errors are often masked as fluent and plausible yet incorrect narratives, impeding timely intervention. Through a longitudinal analysis of a personal assistant LLM agent continuously operating since March 2026 over an eight-week period, the authors conduct root-cause investigations of 22 incidents to propose the first taxonomy of five failure mechanisms specific to LLM agents and formally define the phenomenon of “fail-plausible” behavior. Leveraging a production-grade architecture—comprising 40 scheduled tasks, 8 LLM providers, tool governance agents, and a memory layer—alongside 4,286 unit tests, 827 governance checks, and manual retrospective audits, the study reveals that 70% of silent failures were only detectable by users, while retrospective auditing prevented 87% of recurrence but offered no preemptive mitigation. Failures exhibited latency up to 60 days and predominantly originated from inter-component gaps.
This work addresses the unreliability and poor debuggability of large language models’ memory systems in long-horizon reasoning, often caused by information loss or misaligned retrieval. The paper introduces the first error-tracing and attribution framework specifically designed for memory systems: it constructs an executable memory evolution graph to enable fine-grained, operation-level tracking of information flow; proposes an automated attribution algorithm coupled with a newly developed benchmark, MemTraceBench, to analyze memory failure modes; and leverages attribution signals to drive closed-loop prompt optimization. Empirical evaluation demonstrates that this approach improves end-to-end task performance by up to 7.62%, substantially enhancing both the reliability and interpretability of memory systems.