Training-free Behavior Cloning

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes the Behavior Predictive Control (BPC) framework to address the limited interpretability, high update costs, and retrieval distribution mismatches inherent in neural behavior cloning models. Inspired by behavioral systems theory, BPC synthesizes policies without end-to-end training by integrating an action-aware retrieval metric, a Hankel matrix-based action continuation prior, and closed-form single-step residual correction. By preserving raw demonstration data, the method ensures decision traceability while supporting closed-loop, high-frequency control. Experimental results demonstrate that BPC achieves performance comparable to π0.5 while reducing policy fitting time from hours to seconds. Furthermore, it enables real-time control exceeding 75 Hz on Jetson Orin Nano hardware.
📝 Abstract
Neural behavior cloning compresses demonstrations into large models, making individual actions difficult to trace and policy updates costly. Retrieval policies retain access to demonstrations but struggle with mismatch between recorded and live behavior. We introduce Behavior Predictive Control (BPC), which synthesizes policies without end-to-end policy training by combining an action-aware retrieval metric, a Hankel-based action-continuation prior, and a closed-form one-step residual correction. Inspired by behavioral systems theory, BPC predicts future actions by blending stored observation-action data that best reconstructs the recent runtime observation--action history. Across simulated benchmarks and real-robot deployments, BPC is competitive with learned policies such as $π_{0.5}$ (surpassing it in some cases), while reducing policy fitting from hours to seconds on consumer GPUs and supporting closed-loop control upwards of 75 Hz on a Jetson Orin Nano. The retrieved demonstration windows and their coefficients also provide an intrinsic estimate of task progress. Retaining demonstrations within the deployed policy makes its predictions traceable to supporting trajectories and enables behavior revision through the demonstration bank.
Problem

Research questions and friction points this paper is trying to address.

behavior cloning
retrieval-based policy
training-free imitation learning
policy traceability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Behavior Predictive Control
Training-free
Retrieval-based policy
Hankel matrix
Behavioral systems theory
🔎 Similar Papers