Score
Designs, implements, and evaluates decoders and inference procedures that produce multi-dimensional action vectors simultaneously (vector-wise) instead of generating scalars autoregressively, including model architectures, batching/scheduling strategies, and correctness checks. The work focuses on increasing control-inference throughput while preserving trajectory or action-sequence accuracy and stability under parallel generation.
This work addresses the high rollout latency of autoregressive policies caused by synchronous inference, which hinders their applicability in real-time control requiring low latency and high responsiveness. The authors propose an asynchronous inference framework based on temporal tokenization and constrained decoding, enabling strict latency control and parallel multi-trajectory decoding for the first time in autoregressive policies. While preserving action smoothness, the method significantly improves response speed without sacrificing the fast convergence and strong generalization inherent to autoregressive approaches. Evaluated in both simulation and real-world environments, the proposed framework outperforms state-of-the-art flow-matching policies of comparable scale, achieving substantially higher task completion efficiency while satisfying stringent real-time constraints.
This work addresses the high inference latency and inefficient caching of existing flow-matching–based vision-language-action (VLA) models, which hinder real-time robotic control due to their iterative denoising mechanism. To overcome these limitations, the authors propose an efficient streaming inference framework that exploits the timestep-invariance of perceptual encoders and the denoising process. By partitioning attention context into static, sliding, and dynamic regions, and integrating AdaRMSNorm adaptive normalization, incremental KV cache updates, an asynchronous vision-action decoupled pipeline, and operator fusion, the method achieves significant efficiency gains. Evaluated on the LIBERO and Kinetix benchmarks, it delivers a 2.58× inference speedup, stable 50 Hz output, and up to a 54% reduction in reaction latency—all without compromising task performance.
Robot foundation models often generate actions that violate behavioral correctness and safety constraints, compromising operational safety and logical consistency. Method: We propose the first constraint-decoding framework tailored for robot foundation models, employing Signal Temporal Logic (STL) as a formal specification language to perform real-time verification and correction of action sequences during decoding—without requiring model retraining. The framework supports multimodal inputs, conditional action generation, and dynamic intervention, and is compatible with mainstream navigation foundation models. Contribution/Results: Experiments demonstrate that our approach effectively filters unsafe actions, improves task success rates, and enhances cross-scenario generalization. To the best of our knowledge, this is the first work to enable plug-and-play integration of STL constraints into the inference pipeline of robot foundation models, thereby ensuring safety and logical fidelity in open-world robotic deployment.
This work addresses the inference latency inherent in World Action Models caused by fixed-horizon action chunk generation, which leads to execution pauses, outdated actions, and discontinuous trajectories in robotic systems. To mitigate these issues, the authors propose an asynchronous deployment strategy that overlaps model inference with action execution, thereby enhancing system responsiveness and motion smoothness. Through systematic evaluation of six deployment strategies, the study finds that action blending alone is insufficient to eliminate discontinuities at chunk boundaries. In contrast, prefix-conditioned generation—by precisely aligning observation, prediction, and execution timelines—achieves superior overall performance across dynamic manipulation, precise placement, and long-horizon tasks, striking an optimal balance among task success, execution speed, and trajectory smoothness.
This work addresses the challenge of high computational overhead in existing vision-language-action (VLA) models, which hinders low-latency, high-frequency closed-loop robotic control. The authors propose an asynchronous semantic-action decoupling framework that separates semantic understanding from action generation without modifying the VLA backbone or introducing external planners. In this framework, semantic conditions are updated at a low frequency, while the action module outputs control commands at a high frequency. To mitigate performance degradation caused by semantic lag, the method leverages historical action conditioning and temporally misaligned training. Relying solely on internal VLA interfaces, this minimally invasive approach enables high-frequency control, achieving an action module inference throughput of 35.6 Hz. Experiments on the LIBERO benchmark and real-world robots demonstrate its effectiveness in significantly improving real-time control performance.
Current single-pass inference paradigms constrain the performance of non-deterministic generative models in robotic manipulation. To address this limitation, this work proposes TapSampling—a plug-and-play, inference-time sampling framework that enables policy-agnostic execution refinement by efficiently exploring the action latent space and incorporating a semantically interpretable task-progress prediction verifier. Built upon an action variational autoencoder (Action-VAE), TapSampling is compatible with both diffusion and autoregressive models and requires no fine-tuning. Experiments demonstrate that it significantly enhances task success rates and robustness across diverse general-purpose policies in both simulated and real-world environments.
This work addresses the high inference latency of existing Vision-Language-Action (VLA) models, which stems from temporal redundancy in both visual encoding and diffusion-based policy generation, hindering real-time deployment. The authors propose the first VLA framework that jointly optimizes temporal redundancy across perception and action generation: dynamically updating only visual tokens corresponding to changing image regions at the perception stage, and compressing the diffusion sampling process into an efficient two-step generation at the policy stage, accompanied by an efficiency-oriented training mechanism. Evaluated on Libero, RobotWin, and real robotic platforms, the method achieves over 2× inference speedup while maintaining task success rates as high as 98%.
This work addresses the computational inefficiency of existing vision-language-action (VLA) diffusion models, which rely on multi-step denoising generation and suffer from redundant computation. The authors propose a single-step action generation method that biases the sampling of high-noise timesteps during standard diffusion training, enabling efficient action prediction without requiring teacher models, distillation, or auxiliary objectives. This approach exploits the inherent asymmetry between conditioning inputs and target actions in VLA tasks, demonstrating that merely reshaping the noise distribution during training suffices to match or even surpass the performance of multi-step decoding. Experiments show that the single-step strategy achieves parity with ten-step decoding on the LIBERO benchmark suite; notably, when integrated with a 1.4B vision-language model, it attains a 95.6% success rate on LIBERO-Long and demonstrates strong effectiveness in real-world bimanual robot tasks.