🤖 AI Summary
This study addresses the severe quantization accuracy degradation caused by normalization layers and nonlinear operations such as GELU in Transformers, which impedes pure integer deployment. To overcome this, we propose a unified integer quantization framework that exploits the inherent robustness of Softmax and leverages shared Taylor region primitives to perform division and square root operations entirely within the logarithmic domain. Furthermore, a calibration-free, outlier-aware strategy is designed to optimize LayerNorm, enabling Taylor region reconstruction and full-integer computation under post-training quantization. This work completely eliminates reliance on floating-point hardware, achieving end-to-end pure integer inference. Extensive evaluations on vision and language benchmarks demonstrate an absolute accuracy loss below 1.5%, establishing a new paradigm for efficient edge deployment.
📝 Abstract
Post-training quantization (PTQ) enables efficient deployment, yet transformer architectures remain challenging to quantize due to nonlinear layers. While existing methods attribute accuracy loss to insufficient numerical precision, often necessitating floating-point fallbacks, we demonstrate that degradation is actually driven by specific structural error sources. We find that learned scale parameters in normalization layers and compounded approximations in GELU are the primary error contributors, whereas SoftMax remains inherently robust to aggressive quantization. To address these bottlenecks, we introduce TR-PTQ, a unified integer-only formulation using shared Taylor Region (TR) exponential and logarithm primitives. This approach allows computationally expensive operations, including division and square roots, to be performed entirely in the log-domain via standard integer arithmetic. Combined with a calibration-free, outlier-aware optimization for LayerNorm parameters, our method eliminates the need for floating-point hardware units for nonlinearities, achieving less than 1.5\% absolute accuracy degradation across vision and language benchmarks.