🤖 AI Summary
Edge NPU cross-platform low-bit deployment suffers from opaque and heterogeneous compiler quantization strategies—such as scaling, clipping, and kernel support—leading to significant accuracy fluctuations for the same floating-point model across backends and necessitating repeated manual tuning. To address this, we propose Quant-Trim: a training-time method that jointly integrates progressive pseudo-quantization with backward pruning to suppress outlier-induced scale inflation, thereby enhancing model robustness against diverse hardware quantization schemes. Quant-Trim is hardware-agnostic and supports symmetric/asymmetric, per-tensor/per-channel, and INT8/INT4 configurations without computational graph modification or vendor-specific calibration. Experiments demonstrate that it substantially narrows the accuracy gap between floating-point and low-bit models, reduces reliance on compiler-level tuning, and achieves consistent high performance across platforms in key edge inference metrics—including latency, throughput, energy consumption, and cost.
📝 Abstract
Specialized edge accelerators rely on low-bit quantization, but vendor compilers differ in scaling, clipping, and kernel support, often as black boxes. The same floating-point (FP) checkpoint can therefore yield inconsistent accuracy across backends, forcing practitioners to tweak flags or refactor models to vendor-friendly operator subsets. We introduce Quant-Trim, a training-phase method that produces a hardware-neutral checkpoint robust to backend and precision choices. It combines progressive fake quantization to align training with the deployed integer grid and reverse pruning to tame outlier-driven scale inflation while preserving learnability. Quant-Trim is agnostic to quantization schemes (symmetric/asymmetric,per-tensor/per-channel, INT8/INT4) and requires no vendor-specific graph changes.Across models and tasks, it narrows the FP,low-bit gap, reduces dependence on compiler heuristics/calibration, and avoids per-backend retraining. We report accuracy and edge metrics latency, throughput, energy/inference, and cost under static/dynamic activation scaling and varying operator coverage.