RoboAware: Learning to Coordinate Embodied Skills from Counterfactual Outcomes

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge faced by embodied agents in anticipating which policy family is optimal under current physical states when composing modular skills. To this end, it proposes a hierarchical MDP framework based on the P5 architecture to unify skill semantics, and introduces State-Conditioned Counterfactual Branching (SCB) to compensate for missing observations during training data generation. Furthermore, an Execution-Aware Learning (EAL) mechanism is designed to integrate Monte Carlo Tree Search with Q-learning for value function distillation, enabling dynamic coordination between frozen policies and code generation. Experimental results demonstrate that the proposed approach achieves an overall success rate of 77.0% across 100 tasks and attains state-of-the-art performance of 90.0% on both RoboSuite and RoboTwin benchmarks, significantly outperforming existing baselines.
📝 Abstract
Embodied coding agents can combine modular robot skills with frozen end-to-end policies, yet effective composition requires anticipating which policy family will succeed in the current physical state. We present RoboAware, which builds on coding agents' skill orchestration by learning only a state-conditioned responsibility coordinator from counterfactual outcomes. Inspired by the success of REPL, we propose the $P^5$ schema and formulate a hierarchical MDP based on it. $P^5$ organizes skills uniformly into five semantic stages, defining where responsibility can be compared. To address the lack of counterfactual branch outcomes in existing work, we introduce State-Locked Counterfactual Branching (SCB), which restores the same training state to generate and execute a code block from each admissible family, exposing outcomes that selected-branch experience leaves unobserved. Building on this, we propose Execution-Aware Learning (EAL), which combines Monte Carlo tree search with Q-learning to distill these outcomes into family-conditioned values. At deployment, the coordinator selects the policy family according to observable context, and the frozen coding agent generates the next local code block. Comprehensive single-episode evaluations on 100 tasks show that RoboAware reaches a 77.0% overall success rate, with SOTA averages of 90.0% on RoboSuite, 73.8% on diverse LIBERO-Pro task clusters, and 90.0% on challenging RoboTwin bimanual tasks, outperforming existing code-as-policy and VLA-harness baselines.
Problem

Research questions and friction points this paper is trying to address.

embodied agents
skill composition
policy coordination
counterfactual outcomes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Outcomes
Hierarchical MDP
State-Locked Counterfactual Branching
Execution-Aware Learning
Embodied Coding Agents
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Bohan Zhou
The Chinese University of Hong Kong
X
Xingbei Chen
The Hong Kong University of Science and Technology
E
Emily Huang
Knowin AI
Weilin Ruan
Weilin Ruan
Hong Kong University of Science and Technology (Guangzhou)
Spatio-Temporal Data Mining
H
Haojian Huang
The University of Hong Kong
Y
Yehang Zhang
The Hong Kong University of Science and Technology
Zexi Li
Zexi Li
Alibaba Group
Deep LearningLarge Language ModelsFederated Learning
W
Wenqian Li
Knowin AI
Q
Qize Yu
Peking University
Z
Zetian Song
Peking University
Leyi Wu
Leyi Wu
The Hong Kong University of Science and Technology (Guang Zhou))
Generative Model3D GenerationVideo Generation
J
Jinghao Li
The Chinese University of Hong Kong
M
Mingxuan Song
Peking University
X
Xinrun Xu
University of the Chinese Academy of Sciences
Z
Zongyang Qiu
The Hong Kong University of Science and Technology
Y
Yangkai Wei
Knowin AI
T
Tianyi Zhang
Knowin AI
K
Kaiwen Zhou
Knowin AI
Yinchuan Li
Yinchuan Li
Principal Researcher, Noah's Ark Lab
Generative ModelsEmbodied AIArtificial Intelligence
J
James Cheng
The Chinese University of Hong Kong