🤖 AI Summary
This study addresses the challenge of covariate shift in imitation learning, which causes policies to deviate from training distributions, while online data collection introduces safety risks and noise selection difficulties. We propose the first offline noise calibration method based on generative policy predictive distributions. By leveraging diffusion model properties to estimate injected noise offline and employing partial denoising guidance to resolve error measurement in multimodal action spaces, our approach achieves robust imitation learning without traversing noise levels. Experiments demonstrate that this method surpasses the noise-free DAgger baseline in complex, narrow 3D simulated environments, matching the performance of the post-hoc optimal noise level without requiring additional hyperparameter searches.
📝 Abstract
Policies trained with imitation learning can accumulate errors over time, causing the robot to drift outside the training distribution. Existing methods mitigate this covariate shift by collecting additional data where the policy fails or is likely to fail. The first places the robot in unsafe conditions and the second requires choosing an appropriate noise distribution to collect new expert demonstrations under that noise. We propose Policy-Calibrated DAgger, a method that makes use of the properties of recent generative policies to estimate the policy's noise offline by using its own predicted action distribution. We measure a diffusion policy's spread of predicted actions at observations along the expert trajectory and measure its closed-loop error relative to a recorded trajectory. To address issues with measuring error in a multimodal action space, we guide the policy towards the trajectory during closed-loop control through partial denoising, and use properties of a diffusion model to unnormalize the measured error as if we did not guide it. We experiment in a scenario where a robot is tasked to reach an engine lever in a cluttered and narrow environment and show results in a 3D photorealistic simulator and a 2D planar reacher environment. We show that our method surpasses policies trained with dataset aggregation without noising and matches the performance of the best noise level in hindsight, without requiring a sweep over noise levels.