SAFECAST: Robust Failure Detection for VLA Policies with Contrast-Set Training and Calibration

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the vulnerability of vision-language-action (VLA) policies to distribution shifts—such as cluttered scenes, lighting variations, and novel objects—during deployment, which often leads to failure. Existing failure detection methods exhibit limited generalization due to their reliance on calibration data that must closely match deployment conditions. To overcome this, the paper introduces, for the first time, dual contrastive perturbations in both visual and language modalities to synthesize diverse training and calibration data. Coupled with a latent-state risk probe and functional conformal prediction, this approach substantially enhances the robustness of failure detection under distribution shift. Experiments demonstrate consistent improvements in ROC-AUC across multiple vision-language model backbones in both the DROID real-robot platform and the LIBERO simulation environment, with simulation-to-reality calibration outperforming methods using real data alone.
📝 Abstract
Vision-language-action policies often fail under deployment-time distribution shifts such as clutter, distractor objects, lighting changes, novel objects, altered initial states, and reworded instructions. Hidden-state-based risk probes combined with functional conformal prediction can detect rollout failures, but their reliability depends on calibration data matching deployment conditions. We introduce SAFECAST, which leverages contrast set perturbations to improve hidden-state probe training and calibration for deployment time shift. SAFECAST statistically significantly improves failure detection ROC-AUC scores over a state of the art baseline in both real-world DROID and LIBERO simulation experiments across multiple VLM backbones. We further find that SAFECAST benefits most when both visual and language contrast set perturbations are used to augment data, and that with contrast set perturbations, sim-to-real calibration leads to better probes than using real rollout data only.
Problem

Research questions and friction points this paper is trying to address.

vision-language-action policies
distribution shift
failure detection
calibration
deployment robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

contrast-set training
failure detection
vision-language-action policies
conformal prediction
distribution shift
🔎 Similar Papers
No similar papers found.