🤖 AI Summary
End-to-end learning approaches often suffer from poor generalization and insufficient robustness in high-precision, contact-rich manipulation tasks. To address this limitation, this work proposes a hybrid control framework that leverages offline deep reinforcement learning only during task-critical phases, where data are autonomously and densely collected via an automated mechanism, while relying on conventional motion planning for all other stages. Notably, the method requires neither teleoperation nor online policy updates, substantially improving both learning efficiency and robustness. Evaluated on four real-world tasks, the system achieves an average success rate of 96% using only 2–2.5 hours of autonomously collected data—significantly outperforming the strongest baseline at 55%—and maintains strong performance even in out-of-distribution scenarios.
📝 Abstract
Learned policies trained end-to-end on large datasets often remain brittle in high-precision tasks and struggle with generalization. We find that these limitations largely stem from a lack of structure and focus in data collection. Our key insight is to leverage dense data collection only for the critical segment of contact-rich tasks and to rely on traditional planning during simple free-space motion. We propose an automated data-collection scheme in combination with offline deep reinforcement learning for the critical segment of the task, eliminating reliance on a teleoperator's skill and on online policy updates. Across four challenging real-world tasks, using only 2 to 2.5 hours of autonomous data collection, we achieve an average success rate of 96%, compared to the strongest baseline at 55%. Notably, performance remains high in out-of-distribution scenarios where end-to-end approaches struggle. Our results pave the way for targeted data collection for contact-rich tasks and for high success rates in precision applications.