🤖 AI Summary
This study addresses the challenge of steering pretrained robot policies toward out-of-distribution behaviors via preference learning by proposing the PrefPI framework. This method reformulates preference learning as conditional generative modeling and leverages classifier-free guidance to amplify implicit preference signals, thereby overcoming the limitation of conventional approaches that can only reinforce existing behavioral patterns. Consequently, it enables iterative refinement of both diffusion policies and flow-matching Vision-Language-Action (VLA) models. Experimental results demonstrate that, using merely 150 preference trajectories, the proposed approach significantly alters robot behavior in both simulation and real-world hardware deployments, increasing the object transport height from 10.7 cm to 19.8 cm.
📝 Abstract
We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy. Our key idea is to formulate preference learning as preference-conditioned generative modeling: preferred trajectories define a conditional distribution, whose density ratio with the broader behavior prior provides an implicit preference signal amplified by classifier-free guidance (CFG). Repeating this preference-conditioned modeling and guidance step yields a form of preference-guided policy iteration, turning incremental improvements toward previously inaccessible behaviors. Across diffusion policies and the PI0.5 flow- matching VLA in simulation and the real world, PrefPI produces substantial behavioral shifts with limited feedback. In particular, PrefPI increases object transport height from 10.7 cm to 19.8 cm on real hardware with only 150 preference-labeled trajectories.