🤖 AI Summary
This study addresses the limitation of existing Vision-Language-Action (VLA) model evaluations, which focus solely on task success or failure and thus cannot distinguish between persistent execution under failure and proactive disengagement. To this end, we construct ConflictVLA-Bench upon the LIBERO framework, enabling fine-grained behavioral assessment by contrasting trajectories from conflicting and consistent tasks. Our approach innovatively integrates terminal outcomes with execution dynamics for comprehensive evaluation. Experimental results demonstrate that all tested models exhibit failure persistence, revealing that task failure alone is insufficient to determine whether a model possesses behavioral disengagement capabilities. This work bridges a critical evaluation gap in assessing VLA model responses under invalid premises.
📝 Abstract
While Vision-Language-Action (VLA) models perform strongly on manipulation tasks, their responses to invalid task premises remain underexplored. Existing evaluations of premise conflicts often focus on terminal task outcomes, yet task failure alone cannot distinguish behavioral disengagement from continued pursuit followed by an execution error. We call the latter pattern Failed Persistence. To study this phenomenon, we introduce ConflictVLA-Bench, which pairs conflict rollouts with premise-consistent reference rollouts and evaluates both outcomes and execution processes. Built on LIBERO, the benchmark contains 2,826 prompt-conditioned conflict tasks spanning four conflict families, four structural configurations, and two prompt conditions. Across all eight VLAs, invalid premises reduce original goal completion by at least 17.3 percentage points, with the reduction reaching 56.2 percentage points for OpenVLA. Crucially, even when models succeed on premise-consistent tasks and fail on their matched conflict tasks, they often continue to approach the original targets, retain early trajectory structure, and show limited action magnitude suppression. Failed Persistence therefore recurs across the evaluated models. Explicit premise checking does not consistently produce selective and coordinated behavioral changes. These findings show that terminal failure alone establishes neither behavioral disengagement nor refusal and that outcomes alone are insufficient for VLA evaluation. Experimental data and additional details are available on the project page: https://github.com/EmbodiedAISurvey/ConflictVLA-Bench