Quant-Trim in Practice: Improved Cross-Platform Low-Bit Deployment on Edge NPUs

📅 2025-11-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Edge NPU cross-platform low-bit deployment suffers from opaque and heterogeneous compiler quantization strategies—such as scaling, clipping, and kernel support—leading to significant accuracy fluctuations for the same floating-point model across backends and necessitating repeated manual tuning. To address this, we propose Quant-Trim: a training-time method that jointly integrates progressive pseudo-quantization with backward pruning to suppress outlier-induced scale inflation, thereby enhancing model robustness against diverse hardware quantization schemes. Quant-Trim is hardware-agnostic and supports symmetric/asymmetric, per-tensor/per-channel, and INT8/INT4 configurations without computational graph modification or vendor-specific calibration. Experiments demonstrate that it substantially narrows the accuracy gap between floating-point and low-bit models, reduces reliance on compiler-level tuning, and achieves consistent high performance across platforms in key edge inference metrics—including latency, throughput, energy consumption, and cost.

Technology Category

Machine Learning: Learning on the Edge & Model CompressionComputer Vision: Learning & Optimization for CVKnowledge Representation and Reasoning: Qualitative Reasoning

Application Category

Graph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search engines
📝 Abstract
Specialized edge accelerators rely on low-bit quantization, but vendor compilers differ in scaling, clipping, and kernel support, often as black boxes. The same floating-point (FP) checkpoint can therefore yield inconsistent accuracy across backends, forcing practitioners to tweak flags or refactor models to vendor-friendly operator subsets. We introduce Quant-Trim, a training-phase method that produces a hardware-neutral checkpoint robust to backend and precision choices. It combines progressive fake quantization to align training with the deployed integer grid and reverse pruning to tame outlier-driven scale inflation while preserving learnability. Quant-Trim is agnostic to quantization schemes (symmetric/asymmetric,per-tensor/per-channel, INT8/INT4) and requires no vendor-specific graph changes.Across models and tasks, it narrows the FP,low-bit gap, reduces dependence on compiler heuristics/calibration, and avoids per-backend retraining. We report accuracy and edge metrics latency, throughput, energy/inference, and cost under static/dynamic activation scaling and varying operator coverage.
Problem

Research questions and friction points this paper is trying to address.

Addresses inconsistent accuracy across edge NPUs due to vendor compiler variations
Eliminates need for per-backend model retraining and vendor-specific modifications
Reduces accuracy gap between floating-point and low-bit quantization deployments
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-phase method for hardware-neutral quantization
Combines progressive fake quantization with reverse pruning
Agnostic to quantization schemes without vendor-specific changes
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Rayen Dhahri
Corporate Research and Technology, Carl Zeiss AG
Steffen Urban
Steffen Urban
Carl Zeiss AG
Computer VisionRoboticsPhotogrammetryRemote Sensing