π€ AI Summary
This study addresses the long-horizon unreliability in Vision-Language-Action (VLA) open-loop planning caused by state representation mismatches. We propose a state-space visual reasoning framework built upon the Qwen2.5-VL backbone, which decouples static context from dynamic states. To our knowledge, this work introduces the first recursive latent state modeling mechanism, employing an action-conditioned Gated Recurrent Unit (GRU) to update latent states while leveraging prefix caching for efficient multi-step reasoning. Experimental results demonstrate that the proposed framework significantly outperforms existing methods across multiple benchmarks. Notably, it achieves nearly a 100Γ improvement in decoding speed within Maze scenarios, effectively reducing latency and enhancing overall planning reliability.
π Abstract
Despite rapid progress in vision-language-action (VLA) models, existing reasoning paradigms still face a fundamental \emph{state-representation mismatch} in open-loop planning. Given only an initial observation, models must internally simulate action-conditioned state transitions, whereas text-, pixel-, and latent-space reasoning can suffer from lossy spatial compression, error-accumulating visual generation, and bypass of intermediate latent tokens, respectively, undermining reliable long-horizon planning. We propose \textbf{State-Space Visual Reasoning} (SSVR), which decouples static visual context, language constraints, and a recurrent latent state. SSVR encodes the initial image and instruction once, then conditions each action prediction on the latent state and updates it with an action-conditioned GRU. Using Qwen2.5-VL as the backbone, SSVR achieves 99.5/99.6, 96.3/98.0, and 83.9/90.6 EM/PR on FrozenLake, Maze, and MiniBehavior, substantially outperforming prior methods. Extensive experiments support the effectiveness of recurrent state modeling for VLA open-loop planning across input transformations and transfer settings. By reusing static visual-textual context and updating a compact recurrent state, SSVR supports efficient multi-step inference, achieving up to $98.58\times$ faster Maze decoding rollouts than the evaluated baselines with the prefix cache prebuilt.