ActionUNet: Improving Robustness of VLA Models with Efficient Multi-scale Fine-tuning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of aligning semantic understanding with fine-grained temporal execution in Vision-Language-Action (VLA) models, which compromises robustness in cluttered environments. To this end, we propose an efficient multi-scale fine-tuning framework. Methodologically, a lightweight temporal U-Net is employed to integrate hierarchical priors. Furthermore, we introduce SIREN, a novel continuous action decoding mechanism that combines multi-scale feature fusion with explicit second-order smoothness constraints, thereby ensuring action continuity while effectively suppressing mechanical oscillations. Experimental evaluations on benchmarks such as RoboTwin demonstrate that our approach yields absolute improvements in success rate of up to 11.4% for π0.5, significantly enhancing manipulation robustness and generalization capabilities in real-world scenarios.
📝 Abstract
Vision-Language-Action (VLA) models have shown great promise for robotic manipulation by mapping multi-modal semantics to physical actions. However, this mapping inherently struggles to align these coarse-grained semantics with fine-grained temporal execution. It leaves VLA models with limited generalization and insufficient robustness in cluttered environments. To overcome this issue, we propose ActionUNet, an efficient multi-scale fine-tuning framework that enhances pre-trained VLA models with minimal computational cost. ActionUNet first constructs a lightweight temporal U-Net within the temporal-aligned action feature space to fuse hierarchical structural priors, effectively bridging the scale gap between semantics and temporal executions. Recognizing that multi-scale modeling can disrupt microscopic temporal continuity and cause mechanical oscillations, ActionUNet then employs a conditional SIREN as a continuous action decoder. Equipped with explicit second-order smoothness constraints, this decoder guarantees temporal continuity and reduces high-frequency motion jitter. By smoothing temporal discontinuities from multi-scale fusion, this continuous formulation reduces mechanical execution failures while preserving the base VLA model's generalization and manipulation robustness. Extensive experiments on RoboTwin 2.0 and LIBERO-Plus benchmarks, together with real-world hard evaluations, demonstrate that ActionUNet significantly improves {\pi}0.5 success rates by absolute 9.8%, 6.1%, and 11.4%, respectively, while also generalizing to the regression-based OpenVLA-OFT backbone, highlighting its effectiveness and efficiency as a fine-tuning strategy. Code and implementation details are available at https://github.com/Di-Zhu123/ActionUNet.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
semantic-action alignment
robustness
generalization
robotic manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action models
Multi-scale fine-tuning
Temporal U-Net
Conditional SIREN
Continuous action decoding
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Di Zhu
University of Chinese Academy of Sciences
Z
Ziheng Yan
University of Chinese Academy of Sciences
Fang Wan
Fang Wan
University of Chinese Academy of Sciences
Computer VisionMachine LearningObject DetectionWeakly Supervised Learning