🤖 AI Summary
This study addresses the feasibility-likelihood gap in vision-language-action policies, where locally safe decisions can lead to future infeasibility. We propose VICS-G, a method that derives exact marginal distributions for safe task completion via history-conditioned trajectory laws. By evaluating the quality of feasible futures conditioned on candidates, it establishes selective limited-candidate approximation and optimal-preservation recovery conditions, enabling retraining-free, alert-triggered decision reranking. Evaluated on the Safety-CHORES benchmark, VICS-G reduces cumulative safety costs by 1.9% to 57.5% while limiting success rate degradation to within 2.5 percentage points, effectively balancing safety assurance with task performance.
📝 Abstract
A safe action is not necessarily a viable one. Under a frozen vision-language-action (VLA) policy, an action can be likely and locally admissible yet leave no policy-supported route to safe task completion. We call this the feasibility-likelihood gap: likelihood ranks the current action, whereas feasibility depends on the futures that remain after it. We derive the exact next-block marginal of the history-conditioned policy-environment trajectory law restricted to safe task completion. The derivation exposes a candidate-dependent feasible-future mass with two roles: its support records whether safe completion remains possible under the frozen continuation process, and its magnitude measures how much weighted safe-completion mass is preserved. Exact evaluation is impractical online, so we develop a selective finite-candidate approximation, derive conditions for recovering the best retained viable candidate, and instantiate it as an alarm-triggered, training-free reranker. On Safety-CHORES, VICS-G lowers mean cumulative safety cost by 1.9%-57.5% across six settings while remaining within 2.5 percentage points of policy sampling in success and 0.82 steps in mean episode length. The resulting decoder is tied to an exact policy-relative safe-completion target, yet requires neither retraining of the base policy nor online trajectory rollouts.