The Layer Mystery of VLA: An Information-Theoretical Analysis of VLA Latent Interface

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive computational cost of exhaustive policy sweeping for selecting interface layers in Vision-Language-Action (VLA) model backbones. We propose an efficient layer selection method leveraging an InfoNCE-based proxy metric grounded in information bottleneck theory. Our analysis reveals that most fusion configurations yield suboptimal performance, and we formally derive the equivalence between InfoNCE and action prediction under the information bottleneck framework, establishing it as an optimal criterion for evaluating layer quality in place of conventional trial-and-error approaches. The proposed method reduces computational overhead by 9 to 33 times while decreasing the average selection regret from 17.89% to 3.71%, achieving performance comparable to the best fixed-layer heuristics.
📝 Abstract
Vision-language-action (VLA) policies connect a pretrained vision-language backbone to an action head through a latent interface, but which backbone layers this interface should expose remains unclear. We study single-layer selection and multi-layer fusion for frozen backbones across three pretrained models and two manipulation benchmarks, LIBERO and CALVIN, with three policy-training seeds per configuration. Across three fusion mechanisms and three layer-subset strategies, 47 of 54 configurations underperform the best observed single-layer policy. Our stastical analysis further confirms that fusion's advantage is very limited. However, the best layer varies substantially across backbones and benchmarks, making layer selection consequential and exhaustive policy sweeps expensive. We further derive a reweighting equivalence between the proposed information-bottleneck objectives for action-conditioned InfoNCE and action prediction, motivating InfoNCE as a proxy for layer quality. Empirically, InfoNCE provides the most consistent positive association with policy success among four evaluated proxies. Selecting the layer with the highest InfoNCE score requires 9-33 times less GPU compute than exhaustive policy sweeps and reduces mean selection regret from 17.89 percentage points for deepest-layer selection to 3.71 points across six settings. Its mean regret is close to the 3.28-3.50 points achieved by fixed-layer heuristics optimized retrospectively using all six oracle sweeps, without requiring closed-loop evaluations during selection.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
Layer Selection
Latent Interface
Information Bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action
Information Bottleneck
InfoNCE
Layer Selection
Latent Interface
💼 Related Jobs
No related jobs found.
Y
Yuxiang Liu
University of California, Berkeley
L
Lizhi Yang
California Institute of Technology
Fengze Xie
Fengze Xie
California Institute of Technology
RoboticsRobot LearningControl
A
Aaron Ames
California Institute of Technology
Yisong Yue
Yisong Yue
California Institute of Technology; Asari AI; Latitude AI
Machine LearningArtificial Intelligence