🤖 AI Summary
This study addresses the limitation that existing robot foundation models (RFMs) rely on large-scale training and data while lacking efficient methods to enhance zero-shot performance. To overcome this, it proposes the ARC reasoning recipe, which automatically generates causal action traces from the DROID dataset. By integrating automated annotation with model adaptation techniques, this approach enhances the reasoning capabilities of vision-language-action (VLA) models such as π0.5 and WAM architectures without requiring additional robotic data or retraining. The proposed method establishes a new state-of-the-art on the RoboLab-120 benchmark and improves real-world robot task success rates by 82.2 percentage points, achieving a significant breakthrough in zero-shot control performance at minimal computational cost.
📝 Abstract
The prevailing approach to improving robot foundation models (RFMs) relies on larger models, more robot demonstrations, and costly training at scale. We show that there exists an effective and efficient complementary approach: the right reasoning recipe can substantially improve the zero-shot task performance of existing state-of-the-art RFMs. We refer to this recipe as ARC. It consists of three key ingredients: a reasoning trace, a scalable automatic labeling pipeline, and a strategy for adapting pretrained RFMs to use these traces for control. First, we find that effective reasoning traces should be grounded in the robot's next action and explain its causal structure: why the action is appropriate and what effect it should produce. Second, we show that these traces can be generated automatically from existing demonstrations, enabling us to construct ARC-Trace-DROID from DROID without collecting new robot data. Third, we show how state-of-the-art VLAs such as $π_{0.5}$ and WAMs such as Cosmos3-Nano-Policy can learn to use these traces for control, with fine-tuning and inference tailored to each model's architecture and capabilities. Using ARC, we obtain gains in zero-shot RFM performance that, to our knowledge, are unprecedented without additional robot demonstrations or foundation-scale training. The adapted models establish a new state of the art on RoboLab-120 and MolmoSpaces, with gains of up to 50 percentage points on RoboLab-Reasoning-50. On real robots, ARC improves $π_{0.5}$'s task success by 82.2 percentage points. Project website: https://arc-robot-reasoning.github.io/