🤖 AI Summary
This work addresses the challenge of monitoring robotic task execution when failure modes cannot be exhaustively predefined. We propose a general-purpose, enumeration-free visual anomaly detection method that jointly models camera motion, robot kinematics, and learned optical flow prediction. By fusing observed and predicted optical flow, kinematic constraints, and 3D rigid-body transformation errors, we construct a multi-source anomaly score. A probabilistic U-Net enables uncertainty-aware optical flow prediction, and threshold-based decisioning ensures real-time anomaly detection. Our key contribution is the first unified integration of visual, kinematic, and geometric priors for unsupervised anomaly detection. Evaluated on a book-placing task, our method achieves an AUC of 0.804 and an AP of 0.549—substantially outperforming existing baselines.
📝 Abstract
Execution monitoring is essential for robots to detect and respond to failures. Since it is impossible to enumerate all failures for a given task, we learn from successful executions of the task to detect visual anomalies during runtime. Our method learns to predict the motions that occur during the nominal execution of a task, including camera and robot body motion. A probabilistic U-Net architecture is used to learn to predict optical flow, and the robot’s kinematics and 3D model are used to model camera and body motion. The errors between the observed and predicted motion are used to calculate an anomaly score. We evaluate our method on a dataset of a robot placing a book on a shelf, which includes anomalies such as falling books, camera occlusions, and robot disturbances. We find that modeling camera and body motion, in addition to the learning-based optical flow prediction, results in an improvement of the area under the receiver operating characteristic curve from 0.752 to 0.804, and the area under the precision-recall curve from 0.467 to 0.549.