Catch Me If You Can: Real-Time Feedback Denoising for Responsive VLAs

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究解决了VLA模型在动态环境下的响应性问题,通过结合低频扩散规划和高频视觉反馈的方法,实现了实时动作修正。
📝 Abstract
Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by combining semantic knowledge from pretrained vision-language models with expressive action-generation policies. Diffusion-based action generators are particularly effective for modeling temporally coherent action chunks, but these chunks are typically executed open-loop after inference. This limits responsiveness when objects move, contacts change, or the scene evolves during execution. We propose VLA-Feedback, a two-timescale architecture that combines low-frequency diffusion planning with high-frequency visual feedback. Rather than fully denoising an action chunk before execution, VLA-Feedback retains its final denoising step as a lightweight feedback interface, allowing each action to be corrected using the latest observation before it is executed. This design preserves the expressiveness of the diffusion planner while enabling real-time action correction without rerunning the full vision-language diffusion model. VLA-Feedback matched GR00T on static LIBERO tasks while improving average success on dynamic simulation tasks from 27.5% to 85.0%. On real-robot tasks, it improved average success from 51% to 73%. Additional materials can be found on our project page: https://vla-feedback.github.io.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
real-time feedback
open-loop execution
responsiveness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Real-Time Feedback
Denoising
Two-Timescale Architecture
Diffusion Planning
Y
Yiheng Ji
The University of Texas at Austin
X
Xingru Zhou
The University of Texas at Austin
Luis Sentis
Luis Sentis
Professor of Aerospace Engineering, The University of Texas at Austin
Human-Centered RoboticsRobot Control ArchitecturesWhole-Body ControlHuman-Autonomy Teaming
M
Mingyo Seo
The University of Texas at Austin, University of Central Florida