RE-0: Verified Recursive Improvement of Embodied Code-as-Policy Agents through Local On-Policy Distillation

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of erroneous supervision and training bias in Code-as-Policy agents caused by unreliable teacher models. To this end, we propose a recursive verification strategy improvement framework. Its core innovation lies in introducing the first counterfactual verification-based local intervention mechanism, which selectively adopts only environment-validated teacher corrections. This design enables supervision signals and credit assignment to co-evolve with the student policy, while providing rigorous theoretical guarantees on the lower bound of per-round performance gains. By integrating code-as-policy, online distillation, and recursive optimization techniques, the proposed method substantially enhances both assisted execution and standalone agent performance. Furthermore, it demonstrates strong generalization capabilities across novel robots and scenarios.
📝 Abstract
Code-as-Policy agents accomplish long-horizon embodied tasks by generating and executing code, yet continually improving them with teachers that are stronger but not globally reliable remains a key challenge. Existing distillation methods typically treat the teacher's complete behavior as the supervision target and thus misassign training credit on states where the teacher fails. We propose RE-0, a recursively verified policy improvement framework: rather than assuming that the teacher globally outperforms the student, RE-0 requests local corrections from the teacher on the student's own failure histories and checks in the environment whether each correction is genuinely beneficial; verified corrections yield immediate improvement. Building on this, we propose RE-OPD, which turns verified interventions into supervision for on-policy distillation. Only counterfactually verified teacher interventions provide distribution-level supervision, weighted by their measured local benefit, and the improvement they induce is projected back into the standalone student, so both where supervision is applied and how much credit the teacher receives co-evolve with the student policy. We further prove that the student's per-round gain is lower-bounded by its verified intervention gain up to verification and projection error terms. Experiments across multiple Code-as-Policy embodied tasks show that RE-0 improves both teacher-assisted execution and the standalone student, and generalizes to novel robots and scenes.
Problem

Research questions and friction points this paper is trying to address.

Code-as-Policy agents
embodied tasks
knowledge distillation
policy improvement
credit assignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Code-as-Policy
Recursive Policy Improvement
On-Policy Distillation
Verified Interventions
Embodied Agents
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jiawei Zhang
Jilin University, Changchun, China
Xiangrong Zhang
Xiangrong Zhang
Professor, Xidian University
Image processing and understandingpattern recognitionmachine learning
R
Rui Song
Dalian University of Technology, Dalian, China
H
Huanbin Zhou
Jilin University, Changchun, China
C
Chengye Song
Dalian University of Technology, Dalian, China
H
Hongzhou Wang
Jilin University, Changchun, China