ManiPhysicsBench: Physics-Based Assessment of Object Preservation in VLA Manipulation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of existing Vision-Language-Action (VLA) benchmarks that focus solely on task success rates while neglecting physical object preservation during manipulation. We introduce, for the first time, a distinction between "task success" and "safe success," establishing a physics-based evaluation benchmark for object preservation. Specifically, we construct an asset library by integrating 3D meshes with material properties, employ a physics solver to compute damage thresholds and evaluate contact forces, and introduce continuous grasp labels to mitigate object damage caused by binary supervision. Experiments conducted in LIBERO and SimplerEnv reveal significant safety gaps in mainstream models. Furthermore, although retraining effectively improves safe success rates, it compromises partial generalization capability and overall task completion rates.
📝 Abstract
Vision-language-action (VLA) models aim to perform diverse manipulation tasks, but task success in existing rigid-body benchmarks does not indicate whether they preserve objects. We introduce ManiPhysicsZoo, which consolidates literature-supported material properties, 3D meshes, and supporting references into reusable object assets. Using these assets, a solver-based assessment computes grasp-specific damage thresholds from object geometry, material properties, and recorded grasp conditions and compares them with recorded contact forces to assess potential deformation and fracture. Building on these components, ManiPhysicsBench evaluates object preservation in LIBERO and SimplerEnv across three physics axes and three difficulty levels. Public VLA checkpoints show a substantial gap between task success and safe success, defined as task completion while preserving the object. Their gripper commands concentrate near full opening and closure, with largely similar aggregate distributions across objects, consistent with binary gripper supervision. We examine how object-specific continuous gripper labels change model behavior by retraining a VLA model. The retrained model shows more object-dependent gripping and higher safe success, but lower task success and limited generalization of object-preserving behavior.
Problem

Research questions and friction points this paper is trying to address.

Vision-language-action models
Object preservation
Manipulation benchmarks
Safe success
Physical damage assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action (VLA)
Physics-based assessment
Object preservation
Damage threshold
Continuous gripper control
🔎 Similar Papers
No similar papers found.