🤖 AI Summary
This study addresses the high cost of collecting human demonstrations and the expensive, inefficient invocation of foundation models in robot learning by proposing the CAPEX framework. This method introduces a novel execution-experience-based adaptive demonstration collection mechanism that dynamically adjusts the observation and replanning frequencies of foundation models, efficiently distilling their physical reasoning capabilities into deployable policies such as Diffusion Policy and ACT. Experimental results demonstrate that CAPEX increases the number of successful demonstrations by 4.3 times while reducing per-collection costs by 80%. Furthermore, the resulting policies achieve near-human-level performance, validating the feasibility of leveraging foundation models as a low-cost, scalable data source for robot learning.
📝 Abstract
Robot learning has largely relied on human-teleoperated demonstrations to acquire effective learnable behaviors. However, human-operated data collection processes can be unintuitive, difficult to scale, and inherently asynchronous. We explore an alternative: distilling physical behavior from general-purpose multimodal foundation models into deployable robot policies by using the foundation model itself as an autonomous demonstrator. While sufficiently capable models can generate successful zero-shot manipulation trajectories, repeatedly invoking them during physical execution is slow and expensive, limiting their utility as scalable data generators. As a solution, we introduce CAPEX, an experience-conditioned demonstration collection framework that uses execution experience from previous attempts to adapt how frequently the foundation model must observe, reason, and replan. We evaluate across RoboCasa tasks and on physical Franka and bimanual YAM-arm platforms, measuring task success, model calls, token usage, collection time, and cost. We further train Diffusion Policy and ACT on matched sets of human-teleoperated and foundation-model-generated demonstrations to evaluate the downstream learning value of autonomously collected data. We find that CAPEX increases the number of successful demonstrations by 4.3x while reducing the cost per successful demonstration by 80%. Policies trained on CAPEX-generated data approach the performance of those trained on matched human demonstrations; with longer training, this gap largely closes for policies trained from scratch. These results suggest that foundation models can serve as scalable sources of reusable robot experience. Project page: https://capex-paper.github.io/