EMBER-Bench: Benchmarking Cross-Event Causal Memory in Long-Horizon Embodied Tasks

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing embodied AI benchmarks in evaluating agents' capacity for causal reasoning over historical events during long-horizon tasks. To this end, we introduce the first cross-event causal memory evaluation benchmark, comprising 189 household tasks with fine-grained causal annotations. We propose a bidirectional evaluation mechanism encompassing action prediction and causal backtracking, combined with privileged information ablation experiments to precisely identify model deficiencies. The methodology leverages egocentric videos, meticulous human annotation, and multimodal large language models for evaluation. Experimental results demonstrate that the best-performing model achieves only 61.2% accuracy, substantially lagging behind the 98.3% human baseline, while causal annotations significantly enhance performance, revealing critical cognitive bottlenecks in current models.
📝 Abstract
Lifelong physical agents must reason over extended interactions where past events continue to shape the world long after they disappear from view. Beyond recalling what happened, agents must infer how history changes the current state and constrains future actions. Yet existing embodied and video-memory benchmarks largely focus on historical retrieval and summary, leaving such history-dependent causal reasoning underexplored. We introduce EMBER-Bench, an egocentric benchmark for cross-event causal reasoning in long-horizon embodied tasks, for which we newly created the task design, video recording, and data annotation. It contains 189 household tasks and 699 QA pairs, spanning task progress, failure recovery, external interventions, and compound long-horizon tasks with distant dependencies and prerequisites, with fine-grained event and causal-chain annotations. EMBER-Bench evaluates reasoning in both directions: next-action prediction selects the next action from history, and causal traceback, given that action, identifies the historical event that makes it necessary. Input ablations that add action logs or privileged cause-and-consequence annotations to the video indicate which kind of historical information models fail to use. Among the 16 evaluated models, the highest overall accuracy is 61.2%, compared with a mean of 98.3% across two human evaluators. At paired decision points, correct traceback is not associated with correct next-action prediction. Adding action logs yields a gain of 1.6 points, whereas cause-and-consequence annotations yield an additional gain of 13.0 points on top of that. These results suggest that extracting causal information from past events and converting it into constraints on current actions remains a key difficulty for long-horizon embodied agents. Project Page: https://zhaoalexgoat.github.io/EMBER-Bench/
Problem

Research questions and friction points this paper is trying to address.

embodied AI
causal reasoning
long-horizon tasks
cross-event memory
benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Event Causal Reasoning
Long-Horizon Embodied Tasks
Egocentric Benchmark
Causal Traceback
Input Ablation
🔎 Similar Papers
A
Aoyang Cai
Tsinghua University; Beijing Academy of Artificial Intelligence
B
Boning Zhao
The University of Hong Kong; Beijing Academy of Artificial Intelligence
S
Shaoxuan Xie
Beijing Academy of Artificial Intelligence
D
Dahui Gao
Beijing Academy of Artificial Intelligence
H
Huan Yang
Beijing Academy of Artificial Intelligence
Z
Zhongyuan Wang
Beijing Academy of Artificial Intelligence
Zhiwei Yu
Zhiwei Yu
BAAI
Mutimodality InteractionEmbodied AIKnowledge Based QA/QGComputational Humor
G
Guocai Yao
Beijing Academy of Artificial Intelligence