🤖 AI Summary
This study addresses the challenges of redundant computation and the loss of successful reasoning experiences in Vision-Language-Action (VLA) models during closed-loop control. To this end, we propose FLOWMEM, a framework that transforms successful latent computations into reusable memories. By dynamically retrieving and reassembling latent segments, FLOWMEM constructs coherent reasoning trajectories, which are subsequently refined with current observational evidence to generate actions. The core innovation lies in the dynamic reassembly and refinement of reasoning flows, effectively eliminating computational redundancy. Experimental evaluations demonstrate that FLOWMEM achieves success rates of 48.0% and 77.3% on the RoboMME and LIBERO-Plus benchmarks, respectively, significantly outperforming memory-free baselines.
📝 Abstract
Latent reasoning enables vision-language-action (VLA) models to transform multimodal observations into task-relevant internal states before generating continuous robot actions. While existing methods learn to generate or refine such states for each policy query, they discard successful reasoning after execution and therefore reconstruct similar computation from scratch. We present Reasoning and Flow Memory (FLOWMEM), a unified VLA model that turns successful latent computation into reusable reasoning experience. Rather than appending a fixed retrieved context, FLOWMEM dynamically retrieves and recomposes compatible latent fragments as the embodied context evolves, forming a reasoning route that follows the temporal structure and progress of successful computation. The route is then refined using current visual and proprioceptive evidence before it conditions action generation. Experiments on RoboMME and LIBERO-Plus show that FLOWMEM attains 48.0% and 77.3% success, outperforming memory-free policies by 1.7 and 4.1 percentage points, respectively. These results demonstrate the value of reusing successful latent computation for closed-loop VLA control.