🤖 AI Summary
This study addresses the high inference latency of world action models, which impedes real-time closed-loop robotic control, by proposing a general, training-free acceleration framework. Without requiring architecture-specific designs, this framework synergizes parallel execution with adaptive computation through inter-layer dependency overlapping. Furthermore, it reduces computational overhead by integrating feature caching with Transformer residual reuse strategies. Experimental results demonstrate that the proposed method achieves an 8.9× to 10.7× inference speedup with negligible degradation in success rates. Notably, evaluations on real-world robotic tasks reveal significant performance improvements of 17.2% and 37.2%, respectively, underscoring the framework’s practical effectiveness for enabling responsive robot control.
📝 Abstract
World Action Models (WAMs) combine visual dynamics modeling with action generation, but their high inference latency limits responsive robot control. Recent efforts accelerate inference by removing explicit future-video generation at test time, as in FastWAM, an approach that requires a specially tailored architectural design. More general caching strategies exploit feature redundancy, but redundancy alone does not capture the changing computational demands of closed-loop control. To address these challenges, we present RealtimeWAM, a general, training-free framework that coordinates parallel execution with adaptive computation for low-latency inference across diverse WAM architectures. We exploit layerwise dependencies to overlap observation processing with prediction. However, concurrent branches still compete for GPU resources, limiting the benefit of parallel execution. We therefore adapt computation throughout the pipeline through selective reuse, caching observation features in visually stable regions and reusing Transformer residuals while reserving additional refinement for small predicted adjustments. We evaluate RealtimeWAM on FastWAM and OpenWAM across RoboTwin, LIBERO, and LIBERO-Plus. On an RTX 4090, measured mean inference latencies are 24.09 and 63.09 ms, corresponding to average speedups of 8.90$\times$ and 10.67$\times$. Average success rates are 82.75% and 87.41%, respectively, within 0.02 and 0.53 percentage points of native inference. Across five real-world tasks, RealtimeWAM improves average success rates over native inference by 17.2 and 37.2 percentage points on FastWAM and OpenWAM, respectively.