🤖 AI Summary
This work addresses the significant performance degradation of vision-language-action (VLA) models under low-bit quantization, which stems from the loss of task-relevant evidence. To mitigate this, the authors propose a task-evidence-aware mixed-precision post-training quantization framework that dynamically allocates bit widths by modeling layer-wise sensitivity across temporal execution stages. The approach integrates a task evidence graph with a soft bottleneck objective, leveraging gradient-weighted evidence mapping and metrics for evidence quality and attribution distortion. Bit allocation is optimized under constraints on BitOps and model size. Evaluated on the LIBERO benchmark, the method compresses OpenVLA-OFT from 15.4 GB to 4.1 GB while maintaining an average success rate of 96.3% (compared to 97.1% in BF16) and achieves a 1.52× speedup in inference.
📝 Abstract
We propose Mix-QVLA, a task-evidence-aware mixed-precision PTQ framework for VLA models. Mix-QVLA anchors each quantized variant to the full-precision action-token reference decision and evaluates whether quantization preserves task-relevant evidence across key VLA functional boundaries. It computes normalized gradient-weighted task-evidence maps from boundary activations and compares full-precision and quantized maps using evidence-mass and attribution-distribution distortion, capturing changes in both the strength and allocation of decision-supporting evidence. A soft-bottleneck objective aggregates boundary-level degradation into layer-wise sensitivity scores. Mix-QVLA further models sensitivity throughout task execution, capturing phase-dependent shifts in layer importance rather than assuming a fixed sensitivity profile. The resulting evidence- and time-aware scores guide mixed-precision bit allocation under model-size and BitOps budgets. Extensive evaluations on OpenVLA-style policies show that Mix-QVLA improves the accuracy-efficiency trade-off of low-bit VLA deployment. On LIBERO, Mix-QVLA reduces OpenVLA-OFT memory from 15.4 GB to 4.1 GB, retains 96.3 average success compared with 97.1 for the BF16 model, and achieves a 1.52x inference speedup.