Faster and Better? Benchmark Bugs and Design Limitations Distort the Evaluation of Vision-Language-Action Acceleration

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study reveals spurious performance gains in the accelerated evaluation of Vision-Language-Action (VLA) models caused by benchmark deficiencies. By systematically auditing existing simulation benchmarks through trajectory visualization analysis, success-determinant reconstruction, and physical parameter calibration, we identify implementation loopholes and design limitations. We subsequently propose motion-aware scoring and physical parameter correction schemes to reconstruct evaluation protocols that reflect true model capabilities. Experimental results demonstrate that baseline model rankings are reversed under the corrected protocols, confirming that certain acceleration gains are merely evaluation artifacts. This work establishes a more rigorous and reliable evaluation paradigm for VLA acceleration methods.
📝 Abstract
Simulated manipulation benchmarks are the standard tool for evaluating vision-language-action (VLA) policies and the acceleration methods that reduce their inference latency for on-robot deployment. On these benchmarks, we observe that some training-free acceleration methods, which approximate the baseline policy's computation, achieve higher measured success rates than the baseline itself. Success rates alone cannot establish whether such gains come from better task execution or from evaluation flaws. We therefore investigate two kinds of benchmark flaws behind these gains: bugs, where the implementation does not match the intended task or evaluation protocol, and design limitations, where success criteria and simulation settings do not fully capture how acceleration affects task execution. Starting from tasks with anomalous gains, we localize root causes by plotting object trajectories against checker acceptance regions, and classify the resulting bugs into task consistency, initialization, and reproducibility. Extending this audit to seven benchmarks, including RoboTwin, LIBERO-Plus, and VLABench, we identify 22 bugs of these types and 4 design limitations. For the latter, we revise permissive success checkers, correct unrealistic object masses, and add a motion-aware score that favors smoother actions. Experiments show that bug fixes can reverse method rankings, moving the baseline from last to first on one task. Addressing design limitations can likewise remove anomalous gains: on another task, the baseline moves from 21 percentage points behind an accelerated method to 5 points ahead. Gains attributed to acceleration can therefore be artifacts of the benchmark rather than better task execution. We release our bug fixes and revised benchmark settings to support trustworthy evaluation of VLA acceleration.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
benchmark evaluation
acceleration methods
simulation bugs
design limitations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action
Benchmark Evaluation
Inference Acceleration
Simulation Bugs
Motion-aware Score
🔎 Similar Papers
No similar papers found.
Q
Qiwei Chen
Shanghai Jiao Tong University
K
Kaijun Zhou
Shanghai Jiao Tong University
N
Nuohui Shi
Shanghai Jiao Tong University
Z
Zhiyang Li
Shanghai Jiao Tong University
Y
Yuxuan Feng
Shanghai Jiao Tong University
Jinyu Gu
Jinyu Gu
Shanghai Jiao Tong University
Operating SystemSystem SecurityVirtualization