Reflex: Real-Time VLA Control through Streaming Inference

📅 2026-07-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high inference latency and inefficient caching of existing flow-matching–based vision-language-action (VLA) models, which hinder real-time robotic control due to their iterative denoising mechanism. To overcome these limitations, the authors propose an efficient streaming inference framework that exploits the timestep-invariance of perceptual encoders and the denoising process. By partitioning attention context into static, sliding, and dynamic regions, and integrating AdaRMSNorm adaptive normalization, incremental KV cache updates, an asynchronous vision-action decoupled pipeline, and operator fusion, the method achieves significant efficiency gains. Evaluated on the LIBERO and Kinetix benchmarks, it delivers a 2.58× inference speedup, stable 50 Hz output, and up to a 54% reduction in reaction latency—all without compromising task performance.
📝 Abstract
Flow matching Vision-Language-Action (VLA) models promise precise continuous control, but their iterative denoising nature introduces fundamental incompatibilities with real-time robotics: global timestep injection invalidates KV-caching, forcing a choice between slow $O(N^2)$ re-computation or mathematically incorrect cache reuse. We present \textbf{Reflex}, a framework that enables \textit{real-time streaming inference} for flow matching policies by exploiting the \textit{Timestep-Invariance Property} -- that perception encoders are functionally independent of the denoising loop. Reflex partitions the attention context into static, sliding, and dynamic regions, enabling $O(1)$ incremental cache updates while preserving full-batch-equivalent attention outputs for fixed inputs. To ensure stability under continuous high-frequency inference, we introduce \textit{AdaRMSNorm}, an adaptive normalization layer that prevents BFloat16 numerical collapse by gating on flow phase. We further maximize throughput through an \textit{async pipeline} that decouples visual encoding from action generation, combined with \textit{operator fusion} that reduces kernel overhead. On LIBERO and Kinetix benchmarks, Reflex achieves a 2.58$\times$ inference speedup and 50Hz stable streaming, reducing reaction latency by up to 54\% and enabling efficient deployment without performance degradation.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
real-time control
streaming inference
KV-caching
flow matching
Innovation

Methods, ideas, or system contributions that make the work stand out.

streaming inference
Vision-Language-Action (VLA)
KV-caching
Timestep-Invariance Property
AdaRMSNorm
🔎 Similar Papers
2024-02-022024 IEEE Intelligent Vehicles Symposium (IV)Citations: 1