Is Online Interaction Necessary for Recovery? A Minimalist Approach to Robust Planning via Perturbation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Behavioral cloning is susceptible to covariate shift during closed-loop execution, causing robots to deviate from the demonstration data distribution. This work proposes PARS, a method that achieves robust planning using only existing demonstration data. By perturbing proprioceptive states and anchoring the latter half of trajectories, PARS innovatively provides trajectory-level recovery supervision. Built upon flow models, Transformers, and diffusion models to formulate action sequence policies, this approach differs from conventional input noise augmentation by generating corrective actions without requiring expert queries or environment interaction. Experimental evaluations across 51 RLBench tasks demonstrate that PARS significantly enhances robustness against covariate shift during closed-loop execution.
📝 Abstract
Behavior cloning (BC) is vulnerable to covariate shift during closed-loop execution, where small prediction or execution errors can drive the robot toward states poorly covered by the demonstration data. We focus on action-sequence planning, where a policy predicts a finite-horizon sequence of actions as a reference trajectory for robot execution. Existing approaches to covariate shift often rely on collecting additional corrective demonstrations, requiring further environment interaction and access to an expert or reference policy. We propose Perturbation-Augmented Recovery Supervision (PARS), a simple training approach for improving recovery using only existing demonstrations. PARS perturbs the robot's proprioceptive state and anchors only the later portion of the predicted trajectory to the original demonstration, leaving the earlier portion free to generate corrective motion and recover toward the demonstrated behavior within the planning horizon. Unlike conventional input-noise augmentation, which preserves the original supervision target over the entire prediction horizon, PARS explicitly provides trajectory-level supervision for recovery from perturbed states. PARS requires neither additional environment interaction nor expert queries and can be instantiated across different action-sequence policy classes with only minor modifications to their BC objectives. Experiments on 51 RLBench manipulation tasks with flow-based, transformer-based, and diffusion-based policies demonstrate that PARS improves robustness to covariate shift during closed-loop execution.
Problem

Research questions and friction points this paper is trying to address.

behavior cloning
covariate shift
action-sequence planning
closed-loop execution
robust planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Behavior Cloning
Covariate Shift
Perturbation-Augmented Recovery Supervision
Action-Sequence Planning
Robustness
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Bumgeun Park
Korea Advanced Institute of Science and Technology (KAIST)
Donghwan Lee
Donghwan Lee
KAIST
Decision makingcontroland optimization