FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出FoldQuantVLA框架,通过一致折叠方法实现视觉-语言-动作模型的低比特量化,以减少延迟并保持机器人行为。
📝 Abstract
Low-bit vision-language-action inference must reduce observation-to-action latency while preserving robot behavior. We present FoldQuantVLA, a post-training quantization framework that carries a consistent activation representation through calibration, weight rounding, and native integer execution. It combines channel scaling and block Hadamard transforms with dynamic per-token quantization, without policy retraining. Custom TensorRT plugins execute projections in both the language backbone and iterative action expert with four-bit weights and activations (W4A4) on Ada GPUs and Jetson AGX Orin. Evaluation spans LIBERO, SimplerEnv, and two robot platforms. Across three GR00T checkpoints and $π_{0.5}$, W4A4 achieves $1.20$ to $1.33\times$ speedups over floating-point TensorRT on Orin and $1.25$ to $1.52\times$ on desktop. Retaining language attention-output and feed-forward down projections at eight bits (W8A8) improves held-out action fidelity on all four checkpoints. Across four real-robot tasks, this configuration raises observed GR00T N1.7 success from $80.0\%$ with uniform W4A4 to $92.5\%$ over 80 trials per configuration, with a measured additional Orin latency of 1 ms.
Problem

Research questions and friction points this paper is trying to address.

low-bit quantization
vision-language-action models
latency reduction
robot behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

FoldQuantVLA
low-bit quantization
consistent activation representation
dynamic per-token quantization
TensorRT plugins
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
H
Hung T. Ho
VinRobotics, Vietnam
K
Khanh D. Nguyen
VinRobotics, Vietnam
Q
Quang D. Nguyen
VinRobotics, Vietnam
T
Thanh Q. Duong
VinRobotics, Vietnam
Ngan Le
Ngan Le
University of Arkansas
Artificial IntelligenceMachine LearningComputer Vision
Meng Guo
Meng Guo
Peking University
Task and Motion PlanningMulti-robot SystemsRobotic Manipulation
V
Vien A. Ngo
VinRobotics, Vietnam; Center for AI Research, VinUniversity, Vietnam
A
An T. Le
VinRobotics, Vietnam; Center for AI Research, VinUniversity, Vietnam; Intelligent Autonomous Systems, TU Darmstadt, Germany