π€ AI Summary
Existing embodied agents struggle to effectively accumulate evidence, interpret implicit human preferences, and perform multi-candidate comparisons under partial observability in long-horizon, human-centered tasks. To address this, this work introduces DunphyBench, a benchmark designed to evaluate agentsβ long-term decision-making capabilities under multidimensional human preferences, and proposes MeMento, the first preference-conditioned multimodal memory compression mechanism. MeMento selectively compresses historical multimodal observations based on user preferences and integrates seamlessly into vision-language model (VLM)-based agent architectures. Experimental results demonstrate that MeMento reduces memory overhead by 85.38% while improving decision accuracy by 7.18%, substantially narrowing the performance gap with human-level behavior.
π Abstract
Agents are increasingly expected to act not only as task executors, but also as decision-makers on behalf of human users. This shift requires agents to accumulate evidence over long horizons, interpret implicit user preferences, and compare multiple candidates under partial observations. In this work, we propose DunphyBench, a new benchmark for evaluating agents on long-horizon human-centered embodied decision-making, where the agent must navigate through multiple embodied housing environments and make decisions that align with multi-dimensional human preferences. Unlike standard embodied reasoning tasks that often focus on procedural planning or immediate goal completion, our setting requires agents to integrate multimodal, multi-source input into coherent knowledge that supports complex reasoning across long horizon. The evaluation results reveal that there is a substantial gap between current agents and human performance. Furthermore, our diagnosis of state-of-the-art VLM-driven agents reveals that memory management is one of the bottlenecks, where raw multimodal history introduces noise that hinders decision quality. Motivated by this finding, we design MeMento, a preference-conditioned multimodal memory compressor that selectively compresses decision-relevant information from long-horizon history based on user preferences with a fixed set of memory tokens. Experiments show that MeMento helps VLM-driven agents improve accuracy by 7.18%, while reducing memory usage by 85.38% compared to the strongest baseline.