🤖 AI Summary
This study addresses the high runtime latency and tight coupling between policy reuse and task decision-making inherent in existing "code-as-policy" approaches. We propose RACaP, a framework that shifts code generation to an offline evolutionary phase, deploying frozen, typed policy APIs via a ReAct architecture at runtime to decouple physical mechanism reuse from decision-making. Furthermore, RACaP integrates curriculum learning with rejection sampling fine-tuning to drive autonomous policy evolution, leveraging multimodal memory for failure recovery without source code modification. Evaluated on the LIBERO benchmark, our method significantly outperforms baselines, achieving 45.0% success on LIBERO-PRO while accelerating inference by 13.2× and substantially reducing physical invocation overhead.
📝 Abstract
General-purpose robot agents must learn from experience, transfer to new tasks, and act efficiently. Code as Policies (CaP) methods generate and repair programs at runtime, incurring latency and entangling reusable mechanisms with task-specific decisions. We introduce RACaP, an agentic framework that moves coding to evolution and uses a Reasoning-and-Acting (ReAct) loop to call frozen, typed Policy APIs at deployment. A two-phase strategy combines capability curriculum learning with autonomous self-evolution to improve the APIs, the ReAct harness, and experience memory. The APIs encode reusable physical mechanisms while exposing arguments for runtime adaptation. ReAct combines task-specific working memory, long-term experience memory, and visual feedback to select actions, verify outcomes, and recover from failures without modifying source code. RACaP achieves 54.4% success on LIBERO-90, 45.0% on zero-shot LIBERO-PRO, and 46.0% on LIBERO-Long, compared with at most 4.0% for CaP baselines on long-horizon tasks. On LIBERO-PRO, it achieves 2.5 times the success rate of CaP baselines and a 1.9-fold speedup in median policy time. For efficient on-robot deployment, rejection-sampled fine-tuning distills GPT-5.6 ReAct decisions into Qwen3-VL-8B-Instruct, yielding a 13.2-fold per-decision inference speedup and reducing repeated physical calls from 16 to 4. These results show that separating reusable code from runtime decisions supports continued evolution, effective transfer, and efficient long-horizon control.