Don't Throw Away the Tail: Action Upcycling for Policy Acceleration

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off between efficiency and reactivity in robotic policy execution, where discarding unused actions within action chunks leads to wasted computation. Existing solutions rely on internal model signals or additional sampling. We propose Action Upcycling, a training-free algorithm that exploits the smoothness of discarded actions to adaptively extend the execution horizon until velocity fluctuations trigger termination. This method requires no access to model internals, incurs zero additional sampling cost, is orthogonal to existing acceleration techniques, and remains compatible with various chunking strategies as well as VLA and WAM models. Simulation and real-world manipulation experiments demonstrate that our approach reduces policy inference calls by 1.2× to 1.7× without compromising success rates, significantly lowering computational overhead.
📝 Abstract
Modern robot policies predict a chunk of future actions from a single observation, execute only a prefix, and discard the rest before replanning. Choosing the length of this prefix, the execution horizon, poses a trade-off between reactivity and efficiency. A short horizon keeps the policy reactive to the environment, but requires frequent policy calls. Recent test-time methods adaptively select the horizon for each chunk, but they either read model internals, where the signal must be chosen for each architecture, or draw extra samples, which adds cost. We propose *Action Upcycling*, a training-free algorithm that reuses actions the policy would otherwise discard, without accessing model internals or drawing extra samples. We find that discarded actions stay close to their replanned versions as long as the action velocity remains smooth. Action Upcycling therefore extends the execution horizon up to the point where the velocity begins to fluctuate. Extensive experiments on simulated and real-world manipulation tasks show that Action Upcycling reduces policy calls by 1.2--1.7$\times$ with no loss in success rate, across multiple Vision-Language-Action Models (VLAs) and even a World Action Model (WAM). It applies to any chunked policy at negligible cost and is orthogonal to other policy acceleration methods such as few-step sampling and streaming action decoding, opening a new axis for policy acceleration.
Problem

Research questions and friction points this paper is trying to address.

action chunking
execution horizon
policy acceleration
robot manipulation
Vision-Language-Action Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Action Upcycling
Policy Acceleration
Execution Horizon
Vision-Language-Action Models
Training-free Algorithm
🔎 Similar Papers
No similar papers found.