🤖 AI Summary
This study addresses the challenge of disentangling whether agent decisions during log reuse are driven by outcomes, names, or storage locations. To this end, it proposes ReplayLens, a black-box auditing framework that introduces a novel relation-level auditing paradigm. By employing four single-variable intervention techniques—outcome redistribution, object-pair transfer, consistent renaming, and key-slot reallocation—the method achieves constructive disentanglement of memory dependencies, overcoming the blind spots inherent in traditional endpoint evaluations. Empirical findings demonstrate that score binding constitutes the core mechanism driving decision changes, while also revealing agents' sensitivity to ingestion order. These insights provide theoretical guidance for the secure consolidation and index optimization of agent memory systems.
📝 Abstract
When an agent reuses logged experience, a changed decision may reflect the recorded score, the action's name, or the record's position in storage. Standard memory evaluations do not reveal which relationship drives that change. We introduce ReplayLens, a black-box audit that changes one relationship in the stored history at a time, holds the remaining interface fixed, and measures the resulting decision. Four interventions target four relationships. Outcome reassignment swaps which scores belong to which actions. Pair transport moves intact action-score pairs to new record slots. Consistent renaming relabels actions in both history and menu. Key-slot reassignment changes both score attachment and position. A constructive separation shows why the audit is needed: two memory writers with identical endpoint accuracy respond differently to the same replay, so conventional evaluation cannot resolve the underlying dependence. On black-box LLM interfaces, swapping scores changes decisions while moving intact pairs does not, separating score attachment from record order. A bounded-memory study exposes ingestion-order sensitivity that endpoint comparison misses. In sequential experiment planning, altered historical scores redirect exploration and reduce final utility despite fresh measurements. A code-debugging agent with sealed hidden tests shows the same pattern outside model selection. ReplayLens provides a relationship-level audit for deciding whether logged experience can be merged, reordered, or reindexed safely.