🤖 AI Summary
This study addresses the challenge of diagnosing failures and adaptively adjusting Vision-Language-Action (VLA) models in complex robotic manipulation. We propose a recursive framework distillation method that extracts intervention experiences from a capable agent into reusable guidelines, which are executed by a lightweight agent and iteratively refined via feedback to enable cross-task knowledge accumulation and sharing. A core innovation is the reuse of accumulated knowledge without updating model parameters, overcoming single-agent capability bottlenecks through recursive refinement. Experiments demonstrate that real-world manipulation success rates improve from 37.3% to 64.0%. Furthermore, on the SimplerEnv benchmark, the lightweight agent achieves 66.7%, significantly outperforming both baselines and the unguided capable agent.
📝 Abstract
A central goal in robotics is to enable manipulation across changing tasks and environments. Vision-language-action (VLA) models provide broad manipulation capabilities but can struggle when execution requires diagnosing failures and adapting behavior. Strong agents can discover effective interventions through interaction with these policies. We propose Recursive Harness Distillation to accumulate this experience as reusable guidance across agents. A strong agent distills its experience into a playbook for a light agent, then recursively refines the playbook using the light agent's execution feedback. The resulting playbook enables agents to reuse accumulated intervention knowledge in new task instances without updating model parameters. In real-world manipulation, the harness improves success from 37.3% to 64.0%. On SimplerEnv Bridge, the light agent with the playbook achieves 66.7% success, compared with 41.7% for the GR00T-only baseline, and outperforms the strong agent without a playbook. The same playbook also benefits the strong agent, which reaches 79.2% success. These results demonstrate the feasibility of harness distillation for robotics: intervention experience can be accumulated, refined through execution, and reused across agents to improve manipulation.