Do VLAs Understand and Adapt to the Objects They Handle, or Simply Replay Learned Behaviors?

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether vision-language-action (VLA) models genuinely comprehend the physical properties of objects and adapt their policies accordingly, or merely replay behaviors based on statistical co-occurrences. Employing linear probing and representational similarity analysis (RSA), we systematically evaluate the encoding capacity for physical properties across seven mainstream VLA models and assess their motor adaptability to mass variations within the LIBERO simulation environment. Results indicate that robotic pretraining significantly degrades the linear separability of physical properties, with such features encoded substantially weaker than semantic ones. Furthermore, most VLA models fail to effectively adjust their actions in response to mass changes, leading to diminished task success rates. These findings reveal fundamental limitations in the physical reasoning capabilities of current VLA models.
📝 Abstract
This paper asks whether VLA generalization is grounded in a global understanding of objects'physical properties that enables policies to adapt their motion to unseen setups, or if policies simply replay the motions they've learnt that happen to succeed in new setups. The former reflects genuine generalization; the latter reflects incidental robustness. We first examine awareness of physical properties in seven VLAs by applying linear probing and representational similarity analysis (RSA) to their activations. We find that physical properties, including mass, fragility, deformability, friction and size are less decodable than non-physical properties such as semantic category, material, sound and price in nearly every modality stream. Compared with their base VLMs, robot pre-training weakens the linear encoding of physical properties in the language stream. Neither pre-training nor downstream fine-tuning strengthens the alignment between physical-property differences and activation distances. We then ask whether the weak physical information present in these activations shapes the actions a VLA generates. In a controlled LIBERO case study, we increase the mass of an in-domain object and signal the change through language or vision. Most VLAs use similar lifting behaviour for the heavier and original-mass objects, leading to task success declines. The few exceptions change their behaviour in response to lexical or visual cues rather than to mass itself. These results suggest that VLAs encode physical properties weakly and do not reliably use them to adapt their motion.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
physical properties understanding
generalization
policy adaptation
incidental robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Physical Property Encoding
Representational Similarity Analysis
Linear Probing
Policy Generalization
🔎 Similar Papers
No similar papers found.