🤖 AI Summary
This study addresses the control failure problem in quantized world action models caused by heterogeneous error sensitivity across data streams. To this end, we propose SteerQuant, a framework that introduces an action-guided error steering mechanism and stream-specific activation scaling. These techniques redirect quantization errors toward low-sensitivity computational paths, accommodating heterogeneous data stream requirements without increasing bit-width or duplicating weights. Furthermore, SteerQuant incorporates the Rudder engine to optimize low-bit fused inference. Experimental results demonstrate that this 4-bit quantization scheme incurs less than 0.8% performance degradation on the LIBERO benchmark while achieving a 2.23× inference speedup. Real-world robotic deployment further validates the approach, yielding a 1.35× acceleration with task success rates fully preserved.
📝 Abstract
World-action models (WAMs) jointly generate future world states and actions through iterative denoising, using shared weights to process heterogeneous semantic streams of video, proprioceptive, and action tokens. Quantization reduces inference cost, but comparable numerical errors in different streams can have markedly different effects on final actions, making numerical accuracy alone insufficient for reliable control. We introduce SteerQuant, a 4-bit quantization framework for WAMs that steers errors toward computations with less influence on final actions. It maps how each stream's quantization errors affect final actions and uses this map to guide shared channel scaling. Activation scaling is further calibrated for each stream and denoising step to accommodate changes in activation ranges and action impact. This adapts quantization to different stream requirements without duplicating weights or increasing bit-widths for selected streams. To reduce the extra kernel launches and memory traffic introduced by scaling, we develop Rudder, a 4-bit inference engine for WAMs that fuses scaling and output compensation into low-bit kernels. Under W4A8 and W4A4, SteerQuant maintains mean LIBERO success within 0.8 percentage points of full precision, while delivering up to $2.23\times$ denoising speedup over BF16 across three WAMs with reduced peak GPU memory usage. On a real dual-arm robot, W4A8 deployment achieves a $1.35\times$ end-to-end inference speedup while maintaining average task success relative to BF16.