🤖 AI Summary
This study addresses the prohibitive computational and memory overheads that hinder world action model deployment, as well as the insufficient action generation accuracy caused by existing quantization techniques. To this end, we propose Q-WAM, which introduces the Action Observability Gramian (AOG) metric to quantify the impact of quantization errors and designs an Action Subspace Protection (ASP) mechanism. By incorporating high-precision low-rank branches into critical expert layers, ASP enables 4-bit post-training quantization for both weights and activations while effectively preserving action-sensitive channels at minimal cost. Experimental results demonstrate simulation success rates of 89.6%–93.0% with a 3.1–3.4× memory reduction. Furthermore, real-world robot deployments achieve improvements of 12.8–17.6 percentage points over baselines.
📝 Abstract
World Action Models (WAMs) jointly generate video and robot actions through iterative diffusion and perform strongly in robotic manipulation. However, their prohibitive compute and memory costs pose substantial deployment challenges. Post-training quantization (PTQ) can reduce these costs, but existing PTQ methods such as smoothing and rotation are insufficient to maintain the precision of action generation. To overcome this limitation, we propose Q-WAM, a new 4-bit weight-activation quantization for WAMs that preserves the actions the model generates. Specifically, we introduce the \textit{Action Observability Gramian (AOG)}, which measures how much rounding errors in each weighted combination of a layer's input channels change the final action through all denoising steps. We also develop Action-Subspace Protection (ASP), which keeps the few most action-sensitive channel combinations in a tiny 16-bit low-rank branch and quantizes the complementary weights and activations to 4 bits, both as dense matrix multiplications that run efficiently on GPUs. Finally, to preserve action quality with minimal overhead, we identify the experts that matter most for the generated action by aggregating the AOG-derived action mass across the layers of each expert and apply ASP only to those experts. We evaluate Q-WAM on three WAMs, both in simulation and in real-world deployment. On the RoboTwin 2.0 benchmark, it reaches 89.6--93.0\% average success rate, within 1.1 percentage points of the 16-bit models, while reducing the memory of the quantized blocks by 3.1--3.4$\times$. Our method outperforms the strongest baseline, SVDQuant, by 2.5--8.7 percentage points. On a Unitree G1 humanoid and a bimanual UR3 robot, it improves success over SVDQuant by 12.8-17.6 percentage points.