CogWAM: Aligning Semantic Cognition with World Action Modeling via Event-Driven Interfaces

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient local alignment between semantic reasoning and action prediction in robotic policies by proposing a cognition-guided world action model. Methodologically, it introduces an event-driven persistent semantic state to establish an explicit interface and designs a progress-conditioned attention query mechanism to bridge semantic understanding with physical control, significantly reducing state regeneration overhead. Experimental results demonstrate that the proposed approach achieves a score of 15.56% on the RoboDojo benchmark and attains state-of-the-art performance on the BiCoord task. Furthermore, real-world dual-arm manipulation experiments validate its capability for efficient closed-loop execution.
📝 Abstract
Robot policies increasingly incorporate semantic reasoning and future-world prediction, yet combining these capabilities does not guarantee that local predictions and actions remain aligned with task progress. We introduce CogWAM, a cognition-guided world-action model that establishes an explicit semantic interface between task reasoning and world-action learning through a persistent Semantic State, which stores completed task events and the active subtask. CogWAM updates this state only when observations indicate semantic transitions, allowing task-level context to persist across multiple action chunks. To bridge semantic context with physical prediction and control, CogWAM employs progress-conditioned WORLD and ACTION queries that selectively extract task-relevant information for future-world prediction and action generation. During training, the Semantic State provides shared task-progress context for both branches, while inference removes the future-prediction branch and directly generates actions from observations and the maintained state. We further introduce semantic training strategies to improve transition learning and closed-loop conditioning. Without additional robot-action pretraining, CogWAM achieves 15.56 / 11.70 % Score/SR on RoboDojo and state-of-the-art performance on BiCoord, while real-world experiments demonstrate closed-loop dual-arm manipulation with 16.4 fewer Semantic State regenerations than step-wise updating.
Problem

Research questions and friction points this paper is trying to address.

semantic reasoning
world prediction
robot policy
task alignment
action modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic State
World-Action Model
Event-Driven Interface
Progress-Conditioned Queries
Closed-Loop Manipulation
💼 Related Jobs
No related jobs found.
S
Sen Wang
Xi'an Jiaotong University
L
Liu Liu
Horizon Robotics
X
Xinjiang Wang
Horizon Robotics
Z
Zequn Chen
Horizon Robotics
Haoyi Jiang
Haoyi Jiang
Huazhong University of Science and Technology
Computer VisionAutonomous Driving
T
Taojun Ding
Horizon Robotics
T
Tingyang Xiao
Horizon Robotics
Zhizhong Su
Zhizhong Su
Horizon Robotics
Deep LearningComputer VisionAutonomous DrivingRobotics Learning
J
Jie Wang
University of Science and Technology of China
Sanping Zhou
Sanping Zhou
Xi'an Jiaotong University
Computer VisionMachine Learning