🤖 AI Summary
This study addresses the performance degradation of generalist robot policies in unseen environments and the reliance of existing runtime monitors on task-specific tuning. To overcome these limitations, this work proposes an interactive imitation learning framework grounded in a general progress reward model. By leveraging dense reward signals, the method precisely determines optimal intervention timings. Its gating mechanism is architecture-agnostic, requires no access to internal policy parameters, and generalizes across tasks and environments without retuning. Furthermore, the approach effectively optimizes the trade-off between failure detection accuracy and latency, significantly enhancing the autonomous success rate of downstream policies and enabling efficient human-robot collaboration.
📝 Abstract
Recent advances in robot learning have enabled generalist control policies capable of completing a wide range of tasks. However, their performance degrades when deployed in unseen environments, making it critical to detect failures and teach recovery behaviors. Existing runtime monitoring methods often require task- and policy-specific training or hyperparameter tuning, limiting cross-task deployment and introducing additional overhead during iterative policy updates. We present Reward-DAgger, a robot-gated interactive imitation learning framework that uses dense progress signals from a general-purpose reward model to determine when human intervention is needed. Our approach is agnostic to the underlying policy architecture, requires no access to policy internals, and can be applied across tasks without retuning the gating mechanism. Our results show that Reward-DAgger achieves a better failure-detection accuracy-latency tradeoff than existing runtime monitoring baselines. Across eight simulated and real-world tasks, Reward-DAgger consistently improves the downstream policy's autonomous success rate throughout interactive learning and achieves strong return on human effort, outperforming the baselines in most settings. Importantly, the same gating configuration is used across tasks without task-specific hyperparameter tuning, demonstrating transfer across tasks, environments, and policy architectures. Code and videos are available at https://liralab.usc.edu/reward-dagger.