Wiring Matters: Injection Topology and Initialization of Affordance Heads in Vision-Language-Action Policies

πŸ“… 2026-10-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the interference of auxiliary supervision signals with instruction following in vision-language-action policies. To mitigate this, we propose a stop-gradient residual bridging method that blocks gradient flow to protect the backbone network while injecting intermediate features into action experts via active initialization. Our analysis reveals that initialization serves as the dominant factor governing performance, transforming the bridge from a mere regularization mechanism into a critical load-bearing component. Remarkably, the proposed approach introduces less than 1% additional parameters yet achieves a 96.2% success rate, performing on par with more complex multi-expert architectures and substantially improving instruction-following capabilities.
πŸ“ Abstract
Dense affordance supervision is an appealing auxiliary signal for vision-language-action (VLA) policies, yet naively co-training an affordance head can severely damage instruction following. We present a controlled study of how to wire such a head into a modern VLA on the LIBERO benchmark. Our recipe reads the backbone through a stop-gradient and re-injects an intermediate head feature into the action expert via a learned bridge. The stop-gradient is a precondition: letting affordance gradients reach the backbone drops the policy below the headless base (85.5% vs. 93.1%). With the backbone protected, a same-budget 2*2 ablation over injection topology (concatenation vs. residual) and bridge initialization (zero vs. random) shows initialization is the dominant lever. The best wiring, an actively initialized residual bridge, reaches 96.2%, matching the far more elaborate three-expert AffordanceVLA (95.8%) with under 1% extra parameters. Two probes explain the mechanism: ground-truth affordances fed as an input hurt, and inference-time zeroing shows a lazy bridge acts only as a training-time regularizer while an active bridge becomes load-bearing.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action policies
Affordance supervision
Injection topology
Initialization
Instruction following
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action (VLA)
Affordance Supervision
Stop-Gradient
Injection Topology
Bridge Initialization
Zijian An
Zijian An
Unknown affiliation
Linhan Wang
Linhan Wang
Virginia Tech
Neural RenderingComputer Vision
J
Jiayan Wang
TODO: Jiayan Wang’s affiliation and email
Shijie Geng
Shijie Geng
Senior Applied Scientist, Amazon Store Foundation AI (SFAI)
Multimodal LearningFoundation Models
R
Ran Yang
Virginia Tech, Virginia Seafood Agricultural Research and Extension Center, Hampton, VA 23669, USA
Y
Yiming Feng
Virginia Tech, Virginia Seafood Agricultural Research and Extension Center, Hampton, VA 23669, USA
Lifeng Zhou
Lifeng Zhou
Assistant Professor, Drexel University
Robotics