Score
Design and build a memory-management overlay that tracks and applies explicit state or version information to stored records, retaining superseded entries and labelled transitions so historical and current states can be inspected. Construct query-time evidence packets and resolution logic that return answers conditioned on a requested or inferred state, and analyze the correctness and completeness of state-conditioned retrieval and transition exposition.
This study addresses the critical gap in governance of persistent state—such as memory, credentials, and commitments—in long-running large language model agents, particularly concerning recoverability, auditability, and controlled deprecation. Through a systematic review of 435 publications, the work introduces the first comprehensive model of agent persistent state, integrating multidimensional elements including task logs, credentials, and commitments. Building on this foundation, it proposes AOEP-v0, an evaluation protocol grounded in six dimensions: authoritativeness, scope, mutability, provenance, recoverability, and operability. Distinct from prior approaches that prioritize response quality alone, AOEP-v0 explicitly centers on obligations surrounding state modification and recovery, establishing the first cross-domain governance benchmark for the reliability and controllability of always-on intelligent agents.
Current long-term agent memory systems fail to detect whether stored records are contradictory, outdated, or retracted during retrieval, leading to unreliable outputs. This work proposes the Governed Persistent Memory (GPM) model, which introduces source-bound state semantics and a fail-closed structured release mechanism for the first time. By integrating dual-temporal modeling, five-clause executable constraints, hash-frozen baselines, and end-to-end sealed-service evaluation, GPM enables auditable and controlled memory evolution. Evaluated on a benchmark of 3,600 cases, the model achieves perfect alignment with expected outcomes; the sealed-service evaluation attains 100% accuracy (2,400/2,400), fully correcting all baseline errors without regression, and demonstrates zero inconsistencies across multi-engine differential testing.
Traditional long-term memory systems struggle to model the temporal evolution of user states and are prone to interference from outdated or contradictory information. This work proposes a Temporal Evidence Graph framework that captures state-aware query processing through a hierarchical structure of events, sessions, and topics, enriched with typed temporal, causal, update, and contradiction relations. The approach integrates vector retrieval with graph-based path reasoning, introduces validity annotations to distinguish historical facts from current states, and designs an update-aware seed node selection mechanism coupled with path-grounded evidence generation. Evaluated on long-conversation question answering benchmarks, the method significantly improves performance in temporal and multi-hop reasoning. Ablation studies confirm the critical contributions of the hierarchical structure, update-aware initialization, and path-grounded evidence formulation.
This work addresses the implicit policy adaptation (IPA) gap in personalized agents, which persistently rely on outdated information even after user memory updates. To resolve this, the authors propose StateAuditor, a novel mechanism that retroactively audits draft generations by tracing back through stored states. It leverages large language models to generate candidate transitions between old and new states and employs deterministic code to verify source references and timestamps, triggering repairs only for verifiable transitions. Emphasizing verifiability and temporal consistency over semantic coverage, StateAuditor achieves a single-query VTA score of 0.736 (+5.0) on the STALE benchmark and significantly improves current preference accuracy on HorizonBench (p<0.01). Ablation studies confirm that these gains stem directly from the auditing mechanism itself.
This work addresses the persistent degradation of agent reasoning and tool use caused by erroneous memories—such as contamination, staleness, or misattribution—which existing approaches struggle to correct without discarding valid knowledge. The paper formalizes, for the first time, the post-failure memory recovery problem and introduces a dependency-guided rollback repair mechanism. By constructing a typed memory-action dependency graph, the method tracks downstream effects at runtime, selectively deactivates unreliable memories, and replays only those computations relevant to the final answer. Evaluated on a controlled benchmark of 150 cases, the approach achieves an 85.3% recovery rate—surpassing the best baseline (77.3%)—while fully eliminating error sources and preserving all benign memories. In 50 stress-test scenarios, it attains a 68.0% recovery rate and significantly outperforms baselines, achieving the highest statement invalidation F1 score of 0.669.