Resolving State-Representation Mismatch: State-Space Visual Reasoning for Open-Loop VLA Planning

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the long-horizon unreliability in Vision-Language-Action (VLA) open-loop planning caused by state representation mismatches. We propose a state-space visual reasoning framework built upon the Qwen2.5-VL backbone, which decouples static context from dynamic states. To our knowledge, this work introduces the first recursive latent state modeling mechanism, employing an action-conditioned Gated Recurrent Unit (GRU) to update latent states while leveraging prefix caching for efficient multi-step reasoning. Experimental results demonstrate that the proposed framework significantly outperforms existing methods across multiple benchmarks. Notably, it achieves nearly a 100Γ— improvement in decoding speed within Maze scenarios, effectively reducing latency and enhancing overall planning reliability.
πŸ“ Abstract
Despite rapid progress in vision-language-action (VLA) models, existing reasoning paradigms still face a fundamental \emph{state-representation mismatch} in open-loop planning. Given only an initial observation, models must internally simulate action-conditioned state transitions, whereas text-, pixel-, and latent-space reasoning can suffer from lossy spatial compression, error-accumulating visual generation, and bypass of intermediate latent tokens, respectively, undermining reliable long-horizon planning. We propose \textbf{State-Space Visual Reasoning} (SSVR), which decouples static visual context, language constraints, and a recurrent latent state. SSVR encodes the initial image and instruction once, then conditions each action prediction on the latent state and updates it with an action-conditioned GRU. Using Qwen2.5-VL as the backbone, SSVR achieves 99.5/99.6, 96.3/98.0, and 83.9/90.6 EM/PR on FrozenLake, Maze, and MiniBehavior, substantially outperforming prior methods. Extensive experiments support the effectiveness of recurrent state modeling for VLA open-loop planning across input transformations and transfer settings. By reusing static visual-textual context and updating a compact recurrent state, SSVR supports efficient multi-step inference, achieving up to $98.58\times$ faster Maze decoding rollouts than the evaluated baselines with the prefix cache prebuilt.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action (VLA)
Open-Loop Planning
State-Representation Mismatch
Long-Horizon Planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

State-Space Visual Reasoning
Open-Loop VLA Planning
Recurrent Latent State
State-Representation Mismatch
Efficient Multi-Step Inference
πŸ’Ό Related Jobs
No related jobs found.
J
Junhao Xiao
FDU
H
Haoxiang Zhao
HUST
M
Menghao Fang
TJU
J
Jinkui Zhang
CCNU
J
Jinghan Yu
FDU
X
Xinyu Huang
FDU
Zhiyu Wu
Zhiyu Wu
DeepSeek-AI, εŒ—δΊ¬ε€§ε­¦
MLLMEmotion RecognitionSemi-Supervised Learning
K
Kaiming Xu
FDU
Y
Yi Chen
CCNU
Y
Youjun Bao
Kuaishou
Z
Zhiyuan Ma
HUST