🤖 AI Summary
This study addresses the limitations of world action models arising from fixed-length action chunks, which cause uneven computation allocation, hinder adaptation to complex manipulation, and constrain cloud throughput. To overcome these issues, we propose an adaptive action horizon mechanism based on cubic B-splines that compresses action trajectories into a fixed parameter window, dynamically adjusting temporal resolution and execution span while optimizing training via video-supervised alignment. Furthermore, a Jacobian-Pullback Real-Time Chunking (JP-RTC) technique is introduced to ensure execution continuity during asynchronous deployment. Experiments demonstrate that this framework improves success rates by 8.2% and 4.4% on the LIBERO-Plus and RoboCasa benchmarks, respectively. Additionally, it reduces policy invocation counts by 22%–26% and increases the decoded motion per single invocation on real robots by 1.2×–1.6×.
📝 Abstract
World action models (WAMs) are large embodied policies that jointly predict future video and the actions to execute, emitting a fixed-length action chunk per inference call. Such a policy allocates its computational budget uniformly in time, unable to execute for longer over free-space motion or to spend more inference on contact-rich manipulation, which limits the throughput a WAM can reach when served in the cloud. We present SplineWAM, which adaptively compresses the action trajectory into a fixed-size window of cubic B-spline parameters, fitting the knot times to the characteristics of the motion. One parameter budget then decodes into chunks of varying temporal resolution and duration, and both the executed span and the interval until the next policy call follow from the prediction itself. Aligning the video supervision to the fitted knot times of the demonstration rather than to a uniform grid concentrates the supervised frames where the action trajectory is complex. For asynchronous deployment we introduce Jacobian-Pullback Real-Time Chunking (JP-RTC), which imposes chunk continuity on the decoded raw actions the robot executes rather than on the spline parameters, and corrects the parameters through the decoder so that the executed prefix agrees with the actions already committed. On LIBERO-Plus and RoboCasa, SplineWAM improves success rate over an action chunking WAM by $8.2$ and $4.4$ points while cutting policy calls per episode by 22% and 26%. On three bimanual real-robot tasks under asynchronous execution, it leads or matches the baseline while decoding 1.2 to 1.6 times as much executed motion per call.