Event-Aligned Visual Action Reasoning for World Action Models

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the granularity mismatch in existing world action models, which perform visual predictions at fixed time intervals without distinguishing critical interactions from transitional states. To overcome this limitation, this work proposes an event-aligned visual action reasoning framework. Specifically, it introduces an event-aligned supervision mechanism that organizes visual reasoning around key interaction events to adaptively capture dynamic changes. Furthermore, an execution validity detection head is designed to precisely identify valid prediction segments, thereby eliminating action redundancy inherent in chunk-based reasoning. Experimental results demonstrate that the proposed method improves the success rate by 10.26% on the DOMINO benchmark and achieves strong performance on RoboTwin 2.0. Notably, it also enables successful zero-shot transfer from Level 1 to Level 3 tasks.
📝 Abstract
World-Action Models (WAMs) utilize future visual prediction as an intermediate reasoning process to guide action generation. However, existing WAMs typically structure visual imagination according to predefined temporal intervals, without explicitly accounting for the different roles of task-critical interactions and connecting transitions. We argue that effective visual foresight should align directly with task-relevant interactions and their corresponding reasoning demands. To this end, we introduce an event-aligned visual action reasoning framework that organizes visual-action prediction around interaction events. Through event-aligned visual-action supervision, WAM learns to generate event-aligned visual context in each imagined rollout, placing greater emphasis on critical state changes that inform action generation. This shapes the visual reasoning granularity according to the underlying interaction dynamics, with detailed reasoning around task-critical events and coarser progression through connecting transitions. Furthermore, we introduce an execution validity head that identifies the valid portion of each predicted action sequence, avoiding redundant actions during chunked inference. Experiments demonstrate a 10.26 percentage point improvement in DOMINO success rate over baseline and competitive performance on RoboTwin 2.0. It also transfers from DOMINO Level 1 to Levels 2 and 3 without target-level adaptation.
Problem

Research questions and friction points this paper is trying to address.

World-Action Models
Visual Action Reasoning
Event Alignment
Action Generation
Execution Validity
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Action Models
Event-Aligned Visual Reasoning
Visual-Action Supervision
Execution Validity Head
Chunked Inference
🔎 Similar Papers
No similar papers found.