🤖 AI Summary
To address the high computational and memory overhead hindering the deployment of PETR-style multi-view 3D detection models in autonomous driving, this paper proposes the first full W8A8 quantization scheme tailored to this architecture. Our method introduces three key innovations: (1) optimized LiDAR-ray positional embedding to mitigate modality-scale mismatch between image features and ray embeddings; (2) a Dual-Lookup-Table (Dual-LUT) technique for high-fidelity nonlinear approximation of quantized operators such as Softmax; and (3) a numerically stable pre-Softmax quantization mechanism to preserve activation fidelity. Evaluated across multiple PETR variants, our approach achieves only a 1% mAP drop while reducing inference latency by up to 75%, significantly outperforming existing post-training quantization (PTQ) and quantization-aware training (QAT) methods. This work provides a practical, hardware-efficient pathway for deploying PETR-based 3D detectors on automotive edge platforms.
📝 Abstract
Camera-based multi-view 3D detection is crucial for autonomous driving. PETR and its variants (PETRs) excel in benchmarks but face deployment challenges due to high computational cost and memory footprint. Quantization is an effective technique for compressing deep neural networks by reducing the bit width of weights and activations. However, directly applying existing quantization methods to PETRs leads to severe accuracy degradation. This issue primarily arises from two key challenges: (1) significant magnitude disparity between multi-modal features-specifically, image features and camera-ray positional embeddings (PE), and (2) the inefficiency and approximation error of quantizing non-linear operators, which commonly rely on hardware-unfriendly computations. In this paper, we propose FQ-PETR, a fully quantized framework for PETRs, featuring three key innovations: (1) Quantization-Friendly LiDAR-ray Position Embedding (QFPE): Replacing multi-point sampling with LiDAR-prior-guided single-point sampling and anchor-based embedding eliminates problematic non-linearities (e.g., inverse-sigmoid) and aligns PE scale with image features, preserving accuracy. (2) Dual-Lookup Table (DULUT): This algorithm approximates complex non-linear functions using two cascaded linear LUTs, achieving high fidelity with minimal entries and no specialized hardware. (3) Quantization After Numerical Stabilization (QANS): Performing quantization after softmax numerical stabilization mitigates attention distortion from large inputs. On PETRs (e.g. PETR, StreamPETR, PETRv2, MV2d), FQ-PETR under W8A8 achieves near-floating-point accuracy (1% degradation) while reducing latency by up to 75%, significantly outperforming existing PTQ and QAT baselines.