π€ AI Summary
Deploying the Segment Anything Model (SAM) on resource-constrained devices is hindered by its high computational and memory overhead; existing fixed-bit-width post-training quantization (PTQ) methods suffer from suboptimal accuracy-efficiency trade-offs. Method: This paper proposes the first mixed-precision PTQ framework tailored for SAM, featuring a KL-divergence-based layer importance metric and causal mutual information-guided inter-layer dependency modeling. It further formulates fine-grained bit-width allocation as an integer quadratic programming (IQP) problemβthe first such application in SAM quantization. Results: Under a 4/6-bit mixed-precision configuration, our method achieves up to 20% higher mean average precision (mAP) in instance segmentation and object detection compared to state-of-the-art PTQ baselines, while simultaneously improving model compression ratio and inference latency. The framework establishes a new paradigm for efficient, task-aware SAM deployment.
π Abstract
The Segment Anything Model (SAM) is a popular vision foundation model; however, its high computational and memory demands make deployment on resource-constrained devices challenging. While Post-Training Quantization (PTQ) is a practical approach for reducing computational overhead, existing PTQ methods rely on fixed bit-width quantization, leading to suboptimal accuracy and efficiency. To address this limitation, we propose Mix-QSAM, a mixed-precision PTQ framework for SAM. First, we introduce a layer-wise importance score, derived using Kullback-Leibler (KL) divergence, to quantify each layer's contribution to the model's output. Second, we introduce cross-layer synergy, a novel metric based on causal mutual information, to capture dependencies between adjacent layers. This ensures that highly interdependent layers maintain similar bit-widths, preventing abrupt precision mismatches that degrade feature propagation and numerical stability. Using these metrics, we formulate an Integer Quadratic Programming (IQP) problem to determine optimal bit-width allocation under model size and bit-operation constraints, assigning higher precision to critical layers while minimizing bit-width in less influential layers. Experimental results demonstrate that Mix-QSAM consistently outperforms existing PTQ methods on instance segmentation and object detection tasks, achieving up to 20% higher average precision under 6-bit and 4-bit mixed-precision settings, while maintaining computational efficiency.