The Linear Representation Hypothesis for Vision-Language-Action Models

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of linear representation assumptions in vision-language-action models caused by the dynamic nature of embodied interactions. To overcome this, we propose a signature-based theoretical framework that unifies state representation and policy learning. Methodologically, we construct signature formulations to handle coupled system dynamics, prove that future evolution can be linearly recovered, and introduce a generalized linear model structure incorporating stochastic action chunks. Experiments validate the proposed approach through linear probing, signature-based generalized linear models, and planar control-affine navigation tasks. By constructing explicit oracle representations, this work successfully confirms the predicted linear probing mechanism and demonstrates effective linear guidance within natural parameter spaces, thereby establishing a novel theoretical foundation for embodied intelligence.
📝 Abstract
The linear representation hypothesis (LRH) has become a standard lens for measuring and intervening on semantic information through the internal representations of large language models (LLMs). A growing body of work has begun extending this perspective to vision-language-action (VLA) models, but the dynamical nature of embodied interaction introduces an additional challenge. Unlike semantic attributes commonly studied in LLMs, such as gender or language, a physical quantity of interest (QoI) in a VLA evolves jointly with the system dynamics: the representation influences the actions selected by the policy, which alter the physical state and, in turn, the next representation. In this paper, we develop a theoretical, signature-based formulation of the LRH for VLA that unifies representations and policies. On the representation side, we establish the existence of representations from which the future evolution of a QoI under a candidate action trajectory can be recovered via linear probing. On the policy side, we introduce a signature generalized linear model for stochastic action chunks. This structure yields a monotonic change in the expected future QoI along linear paths in natural parameter space, enabling linear steering. We construct an explicit oracle representation in a planar control-affine navigation experiment and verify the predicted linear probing and steering mechanisms.
Problem

Research questions and friction points this paper is trying to address.

Linear Representation Hypothesis
Vision-Language-Action Models
Embodied Interaction
System Dynamics
Quantity of Interest
Innovation

Methods, ideas, or system contributions that make the work stand out.

Linear Representation Hypothesis
Vision-Language-Action Models
Signature-based Formulation
Linear Steering
Linear Probing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.