🤖 AI Summary
This study addresses the significant gap between high-level decision-making and low-level physical control when deploying large language models in embodied intelligence, where inference latency further constrains real-time operation. We systematically evaluate GPT-6 Astra as a generalist embodied policy across six robotic domains—including grippers, dexterous hands, navigation, and humanoids—under both direct and hybrid control paradigms. Leveraging benchmarks such as RoboDojo, we introduce hybrid architectures integrating π0.5 alongside pretrained whole-body controllers for comprehensive assessment. This work provides the first quantitative characterization of the aforementioned decision-control gap. Results demonstrate that instruction-following navigation achieves a 92% success rate and hybrid control outperforms baselines across multiple tasks. However, direct control exhibits insufficient stability and prohibitive token overhead, revealing critical bottlenecks in transitioning multimodal large models toward end-to-end physical control.
📝 Abstract
GPT-6 Astra exhibits a remarkable ability to generate numerical robot actions, extending its role beyond high-level planning. To assess Astra's capabilities as general-purpose embodied policies, we conduct comprehensive evaluations across six domains, examining direct control, cooperation with learned policies, and feedback-driven adaptation. In gripper manipulation, Astra can correct task targets and prepare contact conditions for subsequent policy execution; hybrid control with π0.5 achieves 48% success on the evaluated RoboDojo subset. In dexterous manipulation, hybrid control achieves 50% success in ten experience-guided DexJoCo trials, while direct in-hand control struggles to coordinate finger contacts. In mobile manipulation, hybrid control reaches 38.7% success on the evaluated RoboCasa365. In navigation, Astra leads our local comparisons, reaching 92% success on RxR instruction following and 82% on HM3D object search, although search incurs substantial detours. In locomotion, dense motion-reference generation remains unreliable: none of five sequential attempts on a single obstacle course reaches the goal, despite improvements in stability and forward progress. In humanoid loco-manipulation, Astra exceeds baseline methods on 13 of 30 HumanoidBench tasks with pretrained whole-body controllers. These findings reveal a gap between useful task decisions and reliable physical control. Inference latency further constrains practical control: across 50 RoboDojo instances per condition, policy-assisted and direct control consume 624.8 million and 1.132 billion tokens. A 30-second locomotion run requires 250 model calls averaging 39.86 seconds each, with physics paused during inference.