🤖 AI Summary
This study addresses the high token overhead in Vision-Language-Action (VLA) models caused by redundant invocations and observations. To mitigate this, we propose PyRUA-Lean, an interactive code execution framework that pioneers encapsulating classical primitives and VLA policies into Python units equipped with conditional checks. By leveraging feedback-driven primitive composition and a selective observation mechanism, the framework returns only essential state information to facilitate efficient replanning. Experimental evaluations in simulated environments such as LIBERO-PRO demonstrate that PyRUA-Lean improves task success rates to 71.7% while reducing large language model invocation counts and input tokens by 49% and 65%, respectively. These results indicate a substantial enhancement in computational efficiency for robotic control.
📝 Abstract
Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant observations incur substantial token overhead. We introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with selective observation: the agent composes classical robot primitives and learned vision-language-action (VLA) policies into Python cells that perform conditional checks and local retries, returning only explicitly requested images and state feedback for replanning. Across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, we compare PyRUA-Lean with a tool-calling baseline using the same GPT-6 Astra planner and underlying robot primitives. Under equal LLM-call budgets, PyRUA-Lean increases overall success from 63.1% to 71.7%. On instances solved by both agents, it uses 49% fewer LLM calls and 65% fewer input tokens.