🤖 AI Summary
Deploying billion-parameter vision–language–action (VLA) models on industrial robots is hindered by the embodiment gap and the high computational cost of full fine-tuning (FFT). This work systematically evaluates low-rank adaptation (LoRA) for fine-tuning the π₀ flow-matching VLA model, conducting four precision assembly tasks on a UR5e robotic arm while exploring varying ranks, parameter allocation schemes, and module-freezing strategies. The experiments demonstrate that LoRA with rank r=32, uniformly applied across the vision encoder, language model, and action expert modules, achieves performance on par with FFT, whereas freezing any core backbone component significantly degrades task success. This approach reduces peak static GPU memory from 36.2 GiB to 10.8 GiB, revealing for the first time that effective embodied adaptation requires concurrent preservation of both semantic and visual plasticity, thereby establishing a practical paradigm for efficient deployment.
📝 Abstract
Deploying billion-parameter Vision-Language-Action (VLA) models on industrial hardware requires fine-tuning to bridge the embodiment gap. Full Fine-Tuning (FFT) provides maximal plasticity but requires data centre-grade GPUs. We present a systematic study of Low-Rank Adaptation (LoRA) for $π_0$, a flow-matching VLA, evaluated on four precision assembly tasks with a UR5e robotic manipulator. Across a sweep of LoRA ranks (r=8 to 256), allocation strategies, and component-freezing ablations, we find no statistically significant advantage of FFT over certain LoRA configurations. Performance saturates at r=32, and uniform allocation across the Vision-Language-Model (VLM) backbone and action expert proves sufficient. Freezing the VLM or restricting the vision encoder to LoRA significantly degrades performance, indicating that embodiment adaptation requires both semantic and visual plasticity. These results suggest that LoRA at r=32 with full vision encoder fine-tuning is a practical approach, reducing static peak VRAM from 36.2 to 10.8 GiB (parameters and optimizer states, activation memory excluded) without detectable performance loss.