π€ AI Summary
This study addresses the interference of auxiliary supervision signals with instruction following in vision-language-action policies. To mitigate this, we propose a stop-gradient residual bridging method that blocks gradient flow to protect the backbone network while injecting intermediate features into action experts via active initialization. Our analysis reveals that initialization serves as the dominant factor governing performance, transforming the bridge from a mere regularization mechanism into a critical load-bearing component. Remarkably, the proposed approach introduces less than 1% additional parameters yet achieves a 96.2% success rate, performing on par with more complex multi-expert architectures and substantially improving instruction-following capabilities.
π Abstract
Dense affordance supervision is an appealing auxiliary signal for vision-language-action (VLA) policies, yet naively co-training an affordance head can severely damage instruction following. We present a controlled study of how to wire such a head into a modern VLA on the LIBERO benchmark. Our recipe reads the backbone through a stop-gradient and re-injects an intermediate head feature into the action expert via a learned bridge. The stop-gradient is a precondition: letting affordance gradients reach the backbone drops the policy below the headless base (85.5% vs. 93.1%). With the backbone protected, a same-budget 2*2 ablation over injection topology (concatenation vs. residual) and bridge initialization (zero vs. random) shows initialization is the dominant lever. The best wiring, an actively initialized residual bridge, reaches 96.2%, matching the far more elaborate three-expert AffordanceVLA (95.8%) with under 1% extra parameters. Two probes explain the mechanism: ground-truth affordances fed as an input hurt, and inference-time zeroing shows a lazy bridge acts only as a training-time regularizer while an active bridge becomes load-bearing.