RoboAware: Learning to Coordinate Embodied Skills from Counterfactual Outcomes
This study addresses the challenge faced by embodied agents in anticipating which policy family is optimal under current physical states when composing modular skills. To this end, it proposes a hierarchical MDP framework based on the P5 architecture to unify skill semantics, and introduces State-Conditioned Counterfactual Branching (SCB) to compensate for missing observations during training data generation. Furthermore, an Execution-Aware Learning (EAL) mechanism is designed to integrate Monte Carlo Tree Search with Q-learning for value function distillation, enabling dynamic coordination between frozen policies and code generation. Experimental results demonstrate that the proposed approach achieves an overall success rate of 77.0% across 100 tasks and attains state-of-the-art performance of 90.0% on both RoboSuite and RoboTwin benchmarks, significantly outperforming existing baselines.