VLA-ZO: Fast Zeroth-Order Adaptation for Vision-Language-Action Models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of distribution shift during the deployment of Vision-Language-Action (VLA) models, alongside the memory constraints and slow convergence associated with zeroth-order optimization. To this end, we propose VLA-ZO, a framework that exploits the computational architecture of VLAs by freezing the vision-language prefix and adapting only the action head. Furthermore, it incorporates conditional state reuse and schedule-aware prefetching to enable efficient zeroth-order adaptation. Experimental results demonstrate that VLA-ZO reduces adaptation time by 25 to 32 times while improving task success rates to 63.58%, thereby achieving efficient and robust online adaptation on resource-constrained platforms.
📝 Abstract
Adapting vision-language-action (VLA) models to deployment-time distribution shifts is important for reliable robotic operation, but conventional first-order adaptation can exceed the memory budget of inference-oriented deployment platforms. Zeroth-order (ZO) optimization offers a forward-only alternative with inference-level memory, but accurate gradient estimation requires many perturbation queries, making naive ZO prohibitively slow for large VLA models. We present VLA-ZO, a framework for fast ZO adaptation that exploits the structure of VLA computation. By confining adaptation to the action side, VLA-ZO keeps the expensive vision-language prefix frozen and reuses its conditioning states across perturbation queries and optimizer steps, while schedule-aware prefetching hides state-transfer overhead. On LIBERO camera-viewpoint shifts, VLA-ZO reduces end-to-end adaptation time by 25.59$\times$ at $q=16$ and 32.54$\times$ at $q=64$ relative to baseline ZO, while improving average task success from 48.27% without adaptation to 58.17% and 63.58%, respectively. These results show that making ZO faster can make larger query budgets practical, providing a promising path toward resource-efficient VLA adaptation on deployment platforms.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action Models
Zeroth-Order Optimization
Distribution Shift
Memory-Efficient Adaptation
Robotic Deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zeroth-Order Optimization
Vision-Language-Action Models
Efficient Adaptation
State Reuse
Schedule-Aware Prefetching
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.