π€ AI Summary
This study addresses the limitation of existing Vision-Language-Action (VLA) models, which lack explicit reasoning over future scene evolution while direct future video generation incurs prohibitive computational costs. To this end, we propose IG-VLA, a framework that introduces the first latent spatiotemporal reasoning mechanism to efficiently βimagineβ future states. By leveraging scene gist memory, the inferred outcomes are internalized into compact tokens, effectively balancing predictive accuracy with deployment efficiency. Experimental results demonstrate that our approach improves the success rate by nearly 6% on the LIBERO-Plus benchmark while reducing inference latency to 169.5 ms, achieving a 6.38Γ speedup. Ultimately, IG-VLA enables robotic manipulation that simultaneously exhibits strong reasoning capabilities and high real-time performance.
π Abstract
Vision-language-action (VLA) models increasingly incorporate intermediate reasoning to improve robotic manipulation, yet existing approaches primarily reason about observed states without explicitly anticipating future scene evolution. Extending such reasoning to explicit future rollouts at every inference step, however, introduces substantial computational overhead. We propose IG-VLA, a VLA reasoning framework that enables models to imagine the future and internalize the gist. Our Latent Spatiotemporal Reasoning learns to imagine task-relevant future scene evolution directly in visual representation space, guiding action prediction without costly pixel-level video generation. To further reduce inference overhead, we introduce Scene Gist Memory, which internalizes reasoning-derived scene-behavior associations into a compact Scene Gist Token, preserving the benefits of future reasoning while bypassing explicit future imagination at inference. Extensive experiments on LIBERO, LIBERO-Plus, and VLABench demonstrate the effectiveness and efficiency of IG-VLA. On the LIBERO-Plus Language suite, both the reasoning and gist policies outperform the strongest baseline by nearly 6% in success rate. The gist policy also achieves up to 6.38x speedup over baselines, reducing inference latency from 1081ms to 169.5ms per action chunk on a single NVIDIA A6000 GPU. These results demonstrate that future spatiotemporal reasoning can be effectively internalized for efficient VLA deployment.