Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing Vision-Language-Action (VLA) model robustness evaluations, which rely solely on task success rates and fail to reveal behavioral degradation under perturbations. We propose a benchmark-agnostic, behavior-level robustness evaluation framework extending the LIBERO/LIBERO-Plus benchmarks. By employing multimodal behavioral feature extraction and statistical variance analysis, our approach quantifies how perturbations affect motion smoothness, execution efficiency, and gripper dynamics within successful trajectories. Our findings demonstrate that models achieving identical success rates exhibit significantly divergent behavioral patterns, confirming that perturbations substantively alter the execution strategies of successful trajectories. This work transcends the constraints of single-metric evaluation by establishing a novel paradigm for VLA robustness assessment that integrates fine-grained behavioral metrics.
📝 Abstract
Vision-Language-Action (VLA) models have achieved high task success rates on robot manipulation task benchmarks. More recently, there has been an emphasis on evaluating the robustness of VLA models to perturbations. However, this robustness is still predominantly measured through Task Success Rate (TSR). In this work, we propose a benchmark-agnostic evaluation framework to measure the behavioural robustness of models by characterising how successful trajectories are executed under perturbation. We implement this methodology by extending the widely-used LIBERO and LIBERO-Plus benchmarks. Across three state-of-the-art VLA models, four LIBERO task suites and seven perturbation conditions, we evaluate changes in both typical successful behaviour and its variability, including metrics of motion smoothness, efficiency and gripper behaviour. We find that perturbations can alter the behaviour of successful trajectories, a phenomenon which cannot necessarily be inferred from TSR alone. Across LIBERO suites, we identify cases where state-of-the-art VLA models achieve comparable TSR under the same perturbation condition, yet behaviour on successful trajectories diverges substantially. Therefore, to have a more robust assessment of task performance, we argue that suitable measures of robustness should capture not only whether a task is completed, but also how the robot behaves while completing it. When evaluating the robustness of VLA models, TSR may be complemented by behavioural evaluation metrics that characterise the nature and variability of successful task execution by robots.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
robustness evaluation
input perturbations
task success rate
behavioral metrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action models
Robustness evaluation
Input perturbations
Behavioral metrics
Task success rate