When to Switch: Reliable Action-Chunk Extension for Vision-Language-Action Models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the robot hesitation and unreliable long-horizon execution caused by expensive Vision-Language-Action (VLA) model inference. To overcome these limitations, this work proposes the RACE framework. Our analysis reveals that prediction errors predominantly occur during sub-skill transitions. Accordingly, we introduce a novel approach that leverages auxiliary one-step denoising to predict switching timings, conditioning action generation on this temporal information to achieve efficient and reliable long-horizon control. Experimental results demonstrate that the proposed method surpasses existing state-of-the-art approaches in simulation. Furthermore, real-world deployment reduces idle time by approximately fivefold while significantly improving task success rates.
📝 Abstract
Vision-Language-Action (VLA) models serve as unified policies for robotic manipulation, yet their expensive inference forces robots to pause between policy calls, resulting in stop-and-go execution that interrupts smooth motion and prolongs task completion. Extending the action chunk reduces policy calls and hence these pauses, but predicting farther into the future makes long-chunk execution unreliable. To understand where this unreliability arises, we analyze action errors within long chunks and find that they concentrate around transitions between manipulation subskills, growing sharply with chunk length. This suggests the importance of transition timing, i.e., when to switch subskills within a chunk. Motivated by this observation, we introduce RACE (Reliable Action-Chunk Extension), a framework that predicts the transition timing from an auxiliary one-step denoising pass and conditions action generation on it. By learning and conditioning on transition timing, RACE reduces errors at subskill transitions and enables reliable execution of longer action chunks. Across simulation benchmarks, RACE outperforms fine-tuning at the same chunk length; with 2x longer chunks, it surpasses recent state-of-the-art and efficient VLAs in success rate, and with 4x longer chunks, it remains competitive. On a real robot, RACE uses 4x longer chunks, which reduces the idle time caused by stop-and-go execution by about 5x, while achieving a higher success rate than fine-tuning with the same chunk length. Code and a real-robot demo are available at https://github.com/Seonghoon-Yu/RACE-VLA
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
action chunk extension
robotic manipulation
subskill transition
stop-and-go execution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Action-Chunk Extension
Transition Timing
Denoising
Robotic Manipulation
🔎 Similar Papers
No similar papers found.