π€ AI Summary
PETR-based methods dominate 3D perception, yet suffer severe performance degradation under INT8 quantization (β58.2% mAP and β36.9% NDS on nuScenes). To address this, we propose a quantization-aware positional encoding transformation: the first quantization-friendly, reparameterizable design of positional embeddings, integrated with per-tensor 8-bit post-training quantization (PTQ) and multi-view geometric modeling. Our method incurs zero additional computational overhead while reducing the accuracy gap between INT8 and FP32 to less than 1% in both mAP and NDSβand even surpasses the original PETRβs FP32 performance. On nuScenes, it improves over the INT8 baseline by +58.2% mAP and +36.9% NDS. Moreover, it is fully compatible with diverse PETR variants, significantly advancing efficient edge deployment of vision-based 3D detection.
π Abstract
PETR-based methods have dominated benchmarks in 3D perception and are increasingly becoming a key component in modern autonomous driving systems. However, their quantization performance significantly degrades when INT8 inference is required, with a degradation of 58.2% in mAP and 36.9% in NDS on the NuScenes dataset. To address this issue, we propose a quantization-aware position embedding transformation for multi-view 3D object detection, termed Q-PETR. Q-PETR offers a quantizationfriendly and deployment-friendly architecture while preserving the original performance of PETR. It substantially narrows the accuracy gap between INT8 and FP32 inference for PETR-series methods. Without bells and whistles, our approach reduces the mAP and NDS drop to within 1% under standard 8-bit per-tensor post-training quantization. Furthermore, our method exceeds the performance of the original PETR in terms of floating-point precision. Extensive experiments across a variety of PETR-series models demonstrate its broad generalization.