🤖 AI Summary
CIM macros suffer from low throughput and high inference error due to physical area constraints and limited ADC precision. To address this, we propose a two-stage model adaptation framework: (1) layer-importance-aware model compression and resource reallocation to maximize CIM array utilization; and (2) quantization-aware training integrated with partial-sum quantization modeling to explicitly compensate for ADC non-idealities. Our approach is the first to enable layer-importance-driven co-optimization of CIM resources and supports concurrent activation of 256 wordlines. Experiments demonstrate a 93% model compression ratio, 90% array utilization, inference accuracy on par with floating-point baselines, and significantly reduced weight loading latency.
📝 Abstract
Computing-in-Memory (CIM) macros have gained popularity for deep learning acceleration due to their highly parallel computation and low power consumption. However, limited macro size and ADC precision introduce throughput and accuracy bottlenecks. This paper proposes a two-stage CIM-aware model adaptation process. The first stage compresses the model and reallocates resources based on layer importance and macro size constraints, reducing model weight loading latency while improving resource utilization and maintaining accuracy. The second stage performs quantization-aware training, incorporating partial sum quantization and ADC precision to mitigate quantization errors in inference. The proposed approach enhances CIM array utilization to 90%, enables concurrent activation of up to 256 word lines, and achieves up to 93% compression, all while preserving accuracy comparable to previous methods.