🤖 AI Summary
This work addresses a critical gap in existing robotic benchmarks, which overlook the challenge of long-horizon tasks requiring agents to continuously track state evolution driven jointly by their own exploration and environmental dynamics. To tackle this, the paper introduces “Task State Horizon” (TSH) as a novel dimension for quantifying task difficulty and presents RoboGraph, a compiler that automatically translates state-transition dependencies into executable symbolic task graphs grounded in spatiotemporal causal relationships—including failures and interventions. The framework enables structured evaluation of agents’ state-tracking capabilities and is accompanied by a benchmark dataset comprising 84 scenarios and 588 episodes. Experiments reveal that 15 state-of-the-art agents exhibit significant performance degradation on high-TSH tasks, exposing fundamental limitations in their ability to maintain, explore, and update task states.
📝 Abstract
Frontier agentic models are increasingly deployed as high-level planners for long-horizon embodied tasks. Existing robotic benchmarks have advanced long-horizon evaluation, but primarily characterize difficulty through action-sequence length and subtask complexity, overlooking a distinct challenge: agents must track evolving task-relevant world states induced by both their exploration and environmental dynamics. We define the span of task-relevant state transitions that an agent must track as task-state horizon (TSH). To evaluate how agent performance varies with TSH, we introduce RoboGraph, a robotic task compiler that translates state-transition dependencies into executable symbolic graphs. Specifically, RoboGraph constructs task-state horizons from spatial and temporal causal dependencies, including those induced by unexpected failures and interventions during task execution. Building on RoboGraph, we release a benchmark comprising 588 episodes across 84 scenes with varying TSHs. Experiments evaluating 15 advanced agentic models in both semantic and visual closed-loop environments show that most models struggle with demanding TSHs, revealing substantial gaps in maintaining, exploring, and updating task-relevant state over long horizon.