🤖 AI Summary
This study addresses the limitation of existing autonomous expert-querying mechanisms in robotics, which overlook how takeover timing affects demonstration quality and policy learning. To this end, we propose TimelyDAgger, a framework that optimizes expert intervention timing by monitoring internal features of Vision-Language-Action (VLA) models and adaptively adjusting decision thresholds. Methodologically, the approach integrates bridging PCA with feedback control to enable dynamic threshold adaptation, and introduces the TASR metric to quantitatively evaluate supervision quality. Experimental results demonstrate that, under an equivalent expert budget, the proposed method achieves competitive failure detection rates while significantly improving policy training success rates across most scenarios.
📝 Abstract
DAgger improves robot policies by aggregating expert supervision from states visited during policy execution. Robot-gated DAgger automates expert queries, allowing the robot to decide when to request expert takeover. While existing gates emphasize detecting the need for assistance, takeover timing also shapes the content of these demonstrations and their value for policy learning. We propose TimelyDAgger, combining Bridge-PCA monitoring of internal vision-language-action (VLA) features with Feedback-guided Threshold Adaptation based on expert behavior to improve takeover timing. We introduce an evaluation framework linking failure detection, takeover timing, and policy improvement, including Target-Aligned Supervision Ratio (TASR) for assessing supervision quality without retraining. Experiments show that takeover timing affects policy learning, with TimelyDAgger achieving competitive failure detection and higher post-training success in most evaluated settings under matched expert-action budgets.