Self-Adaptive VLA for Robust Robot Deployment

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the environmental adaptability degradation of Vision-Language-Action (VLA) models caused by hardware wear or calibration drift, proposing an online adaptation framework that eliminates the need for recalibration. Methodologically, a lightweight context encoder is designed to compress deployment-time visual, proprioceptive, and action signals into latent tokens, which dynamically accommodate environmental shifts via adaptive layer normalization (AdaLN) modulation. Furthermore, a post-training paradigm leveraging self-collected rollout data is introduced, utilizing integrable context tokens to enable iterative self-correction. Experimental results demonstrate that the proposed approach recovers over 80% of performance in bimanual dexterous manipulation tasks, significantly enhancing deployment robustness across diverse workstations.
📝 Abstract
While Vision-Language-Action (VLA) models demonstrate impressive capabilities in robotic manipulation, their memoryless nature renders them brittle to test-time environment shifts, particularly hardware shifts caused by wear or imperfect calibration. Enabling these models to self-adapt during deployment without requiring continuous on-site recalibration remains a critical bottleneck for real-world scalability. In this work, we introduce Self-Adaptive VLA, a novel post-training recipe that enables the policy to iteratively adapt to deployment-time hardware shifts leveraging its own rollouts as context. To do so, we first collect policy rollouts under deliberately injected hardware shifts. We then transform the base policy's training data into shift-conditioned expert demonstrations by pre-compensating the expert actions for these known shifts. Next, we introduce a lightweight, plug-in context encoder that compresses the context, including visual observation, proprioception, and actions in the shifted environment, into a latent context token. This token modulates the policy through adaptive layer normalization (AdaLN). Furthermore, we find that context tokens can be ensembled, allowing the policy to iteratively self-correct and mitigate failures step by step. Extensive experiments across four precision-critical bi-manual and dexterous manipulation tasks show that Self-Adaptive VLA recovers over 80% of the base policy's performance under hardware shifts, such as actuation bias and joint encoder offsets. Moreover, Self-Adaptive VLA enables more robust deployment to new workstations compared to the base policy. Our approach provides a pathway for robust large-scale real-world robot deployments and easier maintenance. See videos at https://icefoxzhx.github.io/self-adaptive-vla.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
hardware shifts
self-adaptation
robot deployment
test-time robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Model
Self-Adaptation
Context Encoder
Adaptive Layer Normalization
Hardware Shift