๐ค AI Summary
This study addresses the challenge of effectively transferring historical memory for embodied agents under varying environmental conditions. Leveraging a shared frozen vision-language model within a simulated warehouse environment, it establishes a benchmark to systematically evaluate the robustness of six memory representations and a working memory baseline against variations in starting poses, path availability, and historical relevance. This work provides the first quantification of the independent sensitivity of different memory representations to experiential mismatches, revealing that robustness along a single dimension does not guarantee overall generalization. Furthermore, it demonstrates that episodic memory is susceptible to interference while summary memory exhibits low retention during path blockages, thereby underscoring the necessity of jointly evaluating both the storage and utilization mechanisms of memory.
๐ Abstract
Memory lets an embodied agent reuse past experience, yet retaining useful information does not ensure that the agent can apply it when conditions change. We present MemTransfer, a benchmark comparing six memory representations, a working-memory baseline and five representations of past experience, under a shared frozen vision-language-model policy. It comprises 100 navigation cases across ten task types in a simulated warehouse, with expert demonstrations supplying the history. Three comparisons vary the starting pose, route availability, and amount and task relevance of history. With one demonstration per task, Full-context and Episodic memory reach 95.3% and 100.0% success at the original demonstration start, but lose 48-49 percentage points at a new test start. Summary changes little between these two test starts, yet with four demonstrations per task it retains a smaller fraction of its unchanged-route success after blocking (39.3%) than Working memory (44.8%) or the two trajectory memories (56-58%). At the new test start, increasing from one to four relevant demonstrations raises Episodic success by 14.3 percentage points, while the other evaluated representations gain no more than 1.3 percentage points. Replacing half of the relevant histories with other-task experience lowers success for both trajectory memories. These results show that robustness to one kind of mismatch does not imply robustness to another, motivating evaluation of both stored information and its use at decision time.