Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection

๐Ÿ“… 2026-08-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the tight coupling between fast and slow modules in existing vision-language-action (VLA) systems, which limits flexibility and efficiency. The authors propose a fully decoupled dual-system adaptive inference framework that dynamically selects between a large-scale pretrained planning model and a lightweight reactive controller via a reinforcement learningโ€“driven, environment-aware switching mechanism. This approach achieves both high task success rates and high-frequency closed-loop control. The architecture supports plug-and-play deployment and enables real-time inference at 93.4 Hz. On the LIBERO benchmark, it matches the task success rate of large-scale baselines while significantly accelerating execution in real-world bimanual manipulation tasks without compromising robustness.
๐Ÿ“ Abstract
Embodied intelligence demands both long-horizon reasoning and real-time closed-loop responsiveness. Recent dual-system Vision-Language-Action (VLA) architectures combine fast reactive control with slow deliberative reasoning to balance inference speed and task success rate. However, existing dual-process VLAs tightly couple the fast module to intermediate representations of the slow module, necessitating end-to-end joint training and limiting modularity, extensibility and flexible system switching. In this paper, we propose Environment-aware Model Selection (EMS), an adaptive VLA inference framework that switches between two fully decoupled systems of different scales through environment-aware model selection. The large-scale deliberative system provides globally consistent trajectory planning to ensure task success, while a lightweight reactive system enables high-frequency closed-loop control. A reinforcement-learning-based switching policy dynamically selects which system to invoke based on real-time feedback, enabling sparse use of the slow system and thereby balancing pretrained knowledge utilisation with runtime efficiency. Our design offers three key advantages over prior hierarchical VLA frameworks: (1) a fully decoupled and modular dual-system architecture that supports plug-and-play model replacement; (2) an adaptive, environment-aware switching strategy; (3) high-frequency inference for responsive closed-loop control. We extensively evaluate EMS in both simulation and real-world environments. On the LIBERO benchmark, EMS achieves success rates comparable to the large-scale baseline while increasing the effective action frequency to 93.4 Hz. The framework further demonstrates strong extensibility in real-world dual-arm manipulation tasks, where it accelerates task completion while maintaining robust performance.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
modularity
model coupling
real-time control
embodied intelligence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action
model decoupling
adaptive switching
embodied intelligence
reinforcement learning
๐Ÿ”Ž Similar Papers
Y
Yuewei Sun
ACSLab, Huawei Technologies
L
Lang Qin
ACSLab, Huawei Technologies
Z
Zechuan Tian
ACSLab, Huawei Technologies
Jingwen Li
Jingwen Li
Sichuan Normal University
Learning to OptimizeDeep Reinforcement LearningCombinatorial Optimization Problems
G
Guiqin Wang
ACSLab, Huawei Technologies
Shengzeng Huo
Shengzeng Huo
PhD student of Robotics, The Hong Kong Polytechnic University
robotic manipulation
W
Wenxin Ren
ACSLab, Huawei Technologies
T
Tao Fang
ACSLab, Huawei Technologies
Xiaochen Zhang
Xiaochen Zhang
Beijing Normal University
G
Guanqing Deng
ACSLab, Huawei Technologies
X
Xiang Wang
ACSLab, Huawei Technologies
Xiaowen Dong
Xiaowen Dong
University of Oxford
signal processingmachine learningnetwork sciencecomputational social science
Q
Qinghai Guo
ACSLab, Huawei Technologies
Yuxin Ma
Yuxin Ma
Southern University of Science and Technology
Information VisualizationVisual AnalyticsHuman-Computer Interaction