CAPEX: Efficiently Distilling Foundation Model Behavior into Deployable Robot Policies through Experience-Adaptive Reasoning

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high cost of collecting human demonstrations and the expensive, inefficient invocation of foundation models in robot learning by proposing the CAPEX framework. This method introduces a novel execution-experience-based adaptive demonstration collection mechanism that dynamically adjusts the observation and replanning frequencies of foundation models, efficiently distilling their physical reasoning capabilities into deployable policies such as Diffusion Policy and ACT. Experimental results demonstrate that CAPEX increases the number of successful demonstrations by 4.3 times while reducing per-collection costs by 80%. Furthermore, the resulting policies achieve near-human-level performance, validating the feasibility of leveraging foundation models as a low-cost, scalable data source for robot learning.
📝 Abstract
Robot learning has largely relied on human-teleoperated demonstrations to acquire effective learnable behaviors. However, human-operated data collection processes can be unintuitive, difficult to scale, and inherently asynchronous. We explore an alternative: distilling physical behavior from general-purpose multimodal foundation models into deployable robot policies by using the foundation model itself as an autonomous demonstrator. While sufficiently capable models can generate successful zero-shot manipulation trajectories, repeatedly invoking them during physical execution is slow and expensive, limiting their utility as scalable data generators. As a solution, we introduce CAPEX, an experience-conditioned demonstration collection framework that uses execution experience from previous attempts to adapt how frequently the foundation model must observe, reason, and replan. We evaluate across RoboCasa tasks and on physical Franka and bimanual YAM-arm platforms, measuring task success, model calls, token usage, collection time, and cost. We further train Diffusion Policy and ACT on matched sets of human-teleoperated and foundation-model-generated demonstrations to evaluate the downstream learning value of autonomously collected data. We find that CAPEX increases the number of successful demonstrations by 4.3x while reducing the cost per successful demonstration by 80%. Policies trained on CAPEX-generated data approach the performance of those trained on matched human demonstrations; with longer training, this gap largely closes for policies trained from scratch. These results suggest that foundation models can serve as scalable sources of reusable robot experience. Project page: https://capex-paper.github.io/
Problem

Research questions and friction points this paper is trying to address.

Robot Learning
Foundation Models
Knowledge Distillation
Data Collection
Scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Foundation Model Distillation
Experience-Adaptive Reasoning
Autonomous Demonstration Collection
Robot Policy Learning
Cost-Efficient Inference
S
Shivam Aarya
School of Interactive Computing, Georgia Institute of Technology
Zhang Xi-Jia
Zhang Xi-Jia
Georgia Institute of Technology
Foundation ModelsInterpretabilityExplainabilityHuman-Robot Interaction
Chengyue Huang
Chengyue Huang
Georgia Institute of Technology
Machine LearningDeep LearningRepresentation Learning
J
Junhyun Kim
School of Interactive Computing, Georgia Institute of Technology
H
Huishu Xue
School of Interactive Computing, Georgia Institute of Technology
H
Hrishit Leen
School of Interactive Computing, Georgia Institute of Technology
R
Roman Yakunin
School of Interactive Computing, Georgia Institute of Technology
Animesh Garg
Animesh Garg
Georgia Institute of Technology, University of Toronto
Robotic ManipulationRobot LearningReinforcement LearningMachine LearningComputer Vision
Zsolt Kira
Zsolt Kira
Associate Professor, Georgia Institute of Technology
Machine LearningPerceptionRoboticsArtificial Intelligence