Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high token overhead in Vision-Language-Action (VLA) models caused by redundant invocations and observations. To mitigate this, we propose PyRUA-Lean, an interactive code execution framework that pioneers encapsulating classical primitives and VLA policies into Python units equipped with conditional checks. By leveraging feedback-driven primitive composition and a selective observation mechanism, the framework returns only essential state information to facilitate efficient replanning. Experimental evaluations in simulated environments such as LIBERO-PRO demonstrate that PyRUA-Lean improves task success rates to 71.7% while reducing large language model invocation counts and input tokens by 49% and 65%, respectively. These results indicate a substantial enhancement in computational efficiency for robotic control.
📝 Abstract
Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant observations incur substantial token overhead. We introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with selective observation: the agent composes classical robot primitives and learned vision-language-action (VLA) policies into Python cells that perform conditional checks and local retries, returning only explicitly requested images and state feedback for replanning. Across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, we compare PyRUA-Lean with a tool-calling baseline using the same GPT-6 Astra planner and underlying robot primitives. Under equal LLM-call budgets, PyRUA-Lean increases overall success from 63.1% to 71.7%. On instances solved by both agents, it uses 49% fewer LLM calls and 65% fewer input tokens.
Problem

Research questions and friction points this paper is trying to address.

Vision Language Model
Token Overhead
Robot Agents
Redundant Observations
Model Invocations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision Language Model
Code Execution Framework
Selective Observation
Token Efficiency
Robot Agents
🔎 Similar Papers
No similar papers found.