Learning to Act under Visual Interruptions with Vision-Language-Action Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of maintaining closed-loop control in vision-language-action models during camera stream interruptions. To this end, it introduces MAIL-Bench and MINT. As the first benchmark for visual interruption, MAIL-Bench systematically evaluates the impact of missing inputs on policy performance. The proposed MINT method integrates optical flow extrapolation with an action-conditioned world model to dynamically supply or retract predicted views, ensuring continuous policy execution under visual deprivation. Experimental results demonstrate that this approach significantly improves task success rates for π0.5 and GR00T N1.5 in camera-loss scenarios. Furthermore, the method has been successfully deployed on the physical AgiBot G2 platform, validating its practical effectiveness.
📝 Abstract
Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, but they are typically developed and evaluated with all camera streams available throughout task execution. When a camera stops delivering frames during task execution, the policy must continue acting without access to subsequent observations from the missing view. Despite its practical importance, how such interruptions affect closed-loop manipulation remains insufficiently understood. To investigate this problem, we introduce MAIL-Bench, a benchmark that evaluates visual interruptions with VLA models. By interrupting different cameras at multiple stages of each policy's successful reference trajectory, MAIL-Bench measures how well policies retain their capabilities when visual inputs become unavailable. Building on this benchmark, we propose MINT, which first trains VLA policies to remain functional under missing visual inputs. At inference time, MINT selectively supplements missing observations using optical-flow extrapolation or an action-conditioned world model, and withdraws predicted views when they become unreliable. Experiments on $\pi_{0.5}$ and GR00T N1.5 show that MINT significantly improves task success under camera loss over the original models. Experiments on AgiBot G2 further demonstrate the real-robot deployment under camera loss. The benchmark is available at https://minglejiang.github.io/Mail-Bench/
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action Models
Visual Interruptions
Robotic Manipulation
Camera Loss
Closed-loop Control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Visual Interruptions
MAIL-Bench
MINT
Optical-flow Extrapolation
🔎 Similar Papers