cim design space exploration

Designs and analyzes architectural variants of compute-in-memory (CIM) accelerators by constructing and searching the design space to compare trade-offs such as latency, throughput, area, and energy across different workload mappings and microarchitectural parameters. Builds models or simulations to evaluate performance across configurations, extract design insights, and recommend architectures or parameter settings optimized for target workloads.

cimdesignspaceexploration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the performance limitations of SRAM-based compute-in-memory (CIM) accelerators, which stem from architectural heterogeneity and inadequate mapping strategies that hinder the balance between computation and memory capabilities. To overcome this challenge, the authors propose a hardware-mapping co-optimization methodology that automatically searches for the optimal hardware configuration and dataflow mapping under a fixed area budget. By introducing a generic CIM macro-array abstraction and an accelerator template, the approach enables cross-architecture adaptability. It further expands the optimization space through fine-grained two-level mapping, accelerator-level scheduling, and macro-level tiling techniques. Experimental results demonstrate that, under the same area constraint, the proposed solution achieves 1.58× higher energy efficiency and 2.11× greater throughput compared to state-of-the-art CIM mapping approaches and outperforms existing advanced CIM accelerators.

compute-storage balancehardware accelerationmapping strategy

Matrix multiplication in ML inference faces significant energy-efficiency bottlenecks, exacerbated by data movement overheads in conventional architectures. Method: This paper systematically addresses three core challenges in Compute-in-Memory (CiM) chip-level integration: CiM type selection, activation timing determination, and optimal deployment location across cache hierarchies (L1/L2). We propose the first CiM-aware “What-When-Where” three-dimensional co-design framework, integrating scalable analytical modeling with customized mapping algorithms to maximize weight reuse and minimize data movement. An INT-8 precision CiM prototype is implemented on a tensor-core-like architecture. Contribution/Results: Experiments demonstrate up to 3.4× energy-efficiency improvement and 15.6× throughput gain over baseline accelerators. The framework provides a quantifiable, reusable methodology and empirical benchmarks for practical CiM deployment in AI accelerators.

Deciding where to integrate CiM in memory hierarchy for efficiency.Determining suitable Compute-in-Memory (CiM) types for ML inference.Identifying optimal timing for CiM use in ML workloads.

This work addresses the limitations of existing machine learning workload partitioning approaches for compute-in-memory (CIM) systems, which often overlook critical RRAM constraints—including storage capacity, high write latency, and endurance—and fail to exploit the full potential of CPU–CIM协同 computation. To overcome these challenges, the paper introduces the first unified integer linear programming (ILP) framework that jointly models RRAM physical constraints, parallelism, and heterogeneous resource scheduling to minimize end-to-end inference latency while respecting hardware limitations. By integrating empirical performance profiling with analytical modeling, the proposed method achieves globally optimal workload partitioning and enables design space exploration for CIM accelerators. Experimental results demonstrate significant speedups of 30.9× and 7.3× over CPU-only execution on edge and high-performance CPU platforms, respectively, substantially enhancing inference efficiency in heterogeneous systems.

Computing-in-MemoryHeterogeneous ComputingMachine Learning

Current digital compute-in-memory (CIM) accelerators lack an end-to-end design framework supporting capacity-constrained modeling and hardware-software co-optimization. To address this, we propose the first capacity-aware CIM instruction set architecture (ISA) along with a dedicated compilation flow that jointly incorporates fine-grained data partitioning, parallel scheduling, and architecture-algorithm co-optimization. Our framework unifies a configurable ISA, a domain-specific compiler, and a cycle-accurate CIM simulator—enabling automated mapping and holistic hardware-software co-modeling. It facilitates rapid prototyping across diverse architectural configurations and enables system-level performance and energy-efficiency evaluation on mainstream DNN models. Experimental results demonstrate significantly improved design space exploration efficiency. This work establishes a foundational infrastructure for the efficient development and optimization of digital CIM accelerators.

Insufficient support for digital CIM capacity constraints in frameworksLack of comprehensive tools for digital CIM accelerator developmentNeed for integrated workflow for DNN evaluation on CIM

Scalable and RISC-V Programmable Near-Memory Computing Architectures for Edge Nodes

Jun 20, 2024
MC
Michele Caon
🏛️ Politecnico di Torino | EPFL

To address the energy inefficiency of von Neumann architectures in edge computing and the high integration complexity and poor software support of existing compute-in-memory (CIM) solutions, this paper proposes a software-friendly near-memory computing (NMC) architecture with low integration overhead. We introduce two novel, configurable NMC microarchitectures—NM-Caesar and NM-Carus—that jointly optimize area, performance, and flexibility. To our knowledge, this is the first NMC design supporting native RISC-V programming via custom instruction extensions and synergistic in-memory/near-memory execution. It integrates an 8-bit quantized matrix multiplication engine. Evaluated against an RV32IMC CPU, our system achieves up to 53.9× reduction in execution time and up to 35.6× improvement in energy efficiency. NM-Carus attains a peak energy efficiency of 306.7 GOPS/W—the highest reported for comparable circuits.

Enhancing performance and energy efficiency for next-gen edge computing nodesOvercoming energy inefficiency in edge computing von Neumann architecturesReducing implementation effort and improving flexibility in Compute-In-Memory solutions

Latest Papers

What's happening recently
View more

This work addresses the limitations of SRAM-based compute-in-memory (CIM) accelerators, which suffer from limited on-chip capacity and high off-chip data movement costs when processing large-scale deep neural networks, compounded by the absence of a systematic methodology for dataflow design. To overcome these challenges, the paper introduces AccelCIM, a framework that establishes the first comprehensive dataflow design space encompassing both CIM macro architecture and macro-array organization. It integrates cycle-accurate simulation with post-layout power-performance-area (PPA) co-evaluation to enable rigorous assessment. Through extensive design space exploration, AccelCIM demonstrates its efficacy on representative large language model workloads, offering both theoretical foundations and practical guidance for designing efficient and scalable SRAM CIM accelerators.

Compute-in-MemoryDataflowDNN Accelerator

This work addresses the thermal constraints of in-orbit data centers, where limited radiative heat dissipation restricts the performance of conventional GPUs due to their high heat density and resulting hotspots that necessitate frequency throttling. To overcome this challenge, the authors propose a “radiator-in-the-loop” co-design framework that, for the first time, jointly optimizes radiative cooling capacity with computational architecture energy efficiency. Through thermal simulations and multi-workload evaluations, they demonstrate that compute-in-memory (CIM) architectures exhibit significantly more uniform thermal distribution and higher TOPS/W efficiency compared to GPUs. Experimental results under realistic orbital thermal constraints show that CIM consistently outperforms GPUs across varying thermal budgets, thereby validating its feasibility and superiority as an AI accelerator for space applications.

AI acceleratorscompute-in-memoryradiative cooling

This work addresses the challenge of energy estimation for nested-loop programs on parallel processor arrays, where traditional simulation-based approaches suffer from poor scalability. To overcome this limitation, the paper proposes a symbolic polyhedral energy modeling method that, for the first time, applies symbolic polyhedral analysis to energy estimation of nested loops. By integrating loop transformation theory with array architecture modeling, the approach explicitly captures the impact of mapping and scheduling decisions on energy consumption. Experimental results demonstrate that the method achieves high-accuracy energy predictions across multiple benchmarks, with computational overhead independent of problem size, thereby significantly enhancing the scalability of design space exploration.

energy analysismapping and schedulingnested loop programs

Existing tools lack automated, cache-aware Roofline modeling support across multiple CPU architectures, hindering effective optimization guidance for high-performance computing applications. This work proposes CARM, the first unified and automated modeling framework spanning x86, ARM, and RISC-V architectures. CARM employs assembly-level microbenchmarks to automatically characterize computational throughput and bandwidth across the entire memory hierarchy, and integrates hardware performance counters with dynamic binary instrumentation for fine-grained bottleneck analysis. The framework supports vectorization across multiple instruction set architectures, achieving a maximum deviation of less than 1% in constructed performance roofs across diverse platforms. By delivering high accuracy and broad applicability, CARM significantly fills the tooling gap for architectures such as AMD and RISC-V in this domain.

automatic benchmarkingCache-aware Roofline Modelcross-architecture

This work addresses the challenge of efficiently optimizing deep neural network inference on compute-in-memory (CIM) crossbar arrays, where the high-dimensional, non-convex design space—shaped by model complexity and heterogeneous layer workloads—hinders effective exploration. To tackle this, we introduce, for the first time, a multi-objective Bayesian optimization framework for system-level co-design of CIM architectures, jointly tuning hardware configurations and per-layer network parameters. Our approach efficiently navigates a design space of up to 50 dimensions and ~10²⁷ possible configurations to identify Pareto-optimal trade-offs among accuracy, energy efficiency, and area. Integrating high-dimensional modeling, layer-granular resource allocation, and a CIM-aware simulator, the method achieves 91.72% accuracy on VGG8/CIFAR-10 and 57.2% on VGG16/Tiny-ImageNet-200, while reducing chip area by up to 65.52%, dynamic energy by 52.07%, and read latency by 13.27%, alongside significantly improved memory utilization.

Compute-In-MemoryCrossbar ArchitectureDeep Neural Network

Hot Scholars

DB

David Bol

ECS group, ICTEAM institute, Université catholique de Louvain
SustainabilityCMOS integrated circuitsLow-power designSmart sensors