🤖 AI Summary
This work addresses the significant challenges in failure detection for surgical robot imitation learning, where the scarcity of failure data, highly variable operational dynamics, and the critical trade-off between missed detections and false alarms complicate reliable monitoring. The authors propose FoMo-FD, a novel approach that, for the first time, employs an action-conditioned flow-matching world model to capture normal short-term visual dynamics. Faults are detected at the window level by measuring the inconsistency in inverse transport of latent variables at observation endpoints, entirely without requiring failure examples. Integrating conformal calibration to set task-adaptive thresholds and leveraging multi-view inputs—with emphasis on the wrist-mounted camera—the method achieves a 96.6% detection rate and only 1.3% false positive rate across four surgical tasks and 20 failure modes on the dVRK platform, substantially outperforming existing approaches.
📝 Abstract
Imitation learning has shown increasing promise for autonomous robotic surgery, yet safe deployment remains challenging due to the safety-critical nature of surgical tasks and the complexity and variability of surgical environments. Failure detection is therefore an essential safeguard, but its development remains difficult due to the challenges of scarce failure data, highly variable manipulation dynamics, and the need to balance missed detections against disruptive false alarms. To address these challenges, we introduce FoMo-FD (Flow-Matching World Model for Failure Detection), a failure detection method that learns nominal short-horizon visual dynamics with an action-conditioned flow-matching world model. FoMo-FD scores the inverse-transport nonconformity of observed endpoint latents, enabling window-level detection of visual-action inconsistencies without requiring failure demonstrations. Detection thresholds are obtained by conformal calibration on successful executions, yielding task-specific alarms without assuming future failure types. We evaluate FoMo-FD on four surgically relevant manipulation tasks with twenty failure modes across simulation and real-world experiments using the da Vinci Research Kit (dVRK). Results show that FoMo-FD outperforms observation-level anomaly baselines and a prediction-error variant of the same world model, with the wrist-camera view achieving the strongest performance, including a 96.6% failure detection rate (FDR) at a 1.3% false alarm rate (FAR).