Computing-In-Memory Aware Model Adaption For Edge Devices

📅 2025-10-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
CIM macros suffer from low throughput and high inference error due to physical area constraints and limited ADC precision. To address this, we propose a two-stage model adaptation framework: (1) layer-importance-aware model compression and resource reallocation to maximize CIM array utilization; and (2) quantization-aware training integrated with partial-sum quantization modeling to explicitly compensate for ADC non-idealities. Our approach is the first to enable layer-importance-driven co-optimization of CIM resources and supports concurrent activation of 256 wordlines. Experiments demonstrate a 93% model compression ratio, 90% array utilization, inference accuracy on par with floating-point baselines, and significantly reduced weight loading latency.

Technology Category

Machine Learning: Learning on the Edge & Model CompressionCognitive Modeling & Cognitive Systems: Neural Spike CodingComputer Vision: Learning & Optimization for CV

Application Category

User Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
Computing-in-Memory (CIM) macros have gained popularity for deep learning acceleration due to their highly parallel computation and low power consumption. However, limited macro size and ADC precision introduce throughput and accuracy bottlenecks. This paper proposes a two-stage CIM-aware model adaptation process. The first stage compresses the model and reallocates resources based on layer importance and macro size constraints, reducing model weight loading latency while improving resource utilization and maintaining accuracy. The second stage performs quantization-aware training, incorporating partial sum quantization and ADC precision to mitigate quantization errors in inference. The proposed approach enhances CIM array utilization to 90%, enables concurrent activation of up to 256 word lines, and achieves up to 93% compression, all while preserving accuracy comparable to previous methods.
Problem

Research questions and friction points this paper is trying to address.

Optimizing model compression for CIM macro size constraints
Enhancing quantization-aware training with ADC precision considerations
Improving CIM array utilization while maintaining model accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model compression and resource reallocation for efficiency
Quantization-aware training with partial sum quantization
Enhanced CIM array utilization and concurrent activation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Ming-Han Lin
Institute of Electronics, National Yang Ming Chiao Tung University, Taiwan
T
Tian-Sheuan Chang
Institute of Electronics, National Yang Ming Chiao Tung University, Taiwan