SPIMOE: Exploiting Hybrid Sparsity for Reasoning MoE Inference on Heterogeneous PIM Architectures

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the hardware utilization bottlenecks in long-reasoning Mixture-of-Experts (MoE) models caused by KV cache bloat and expert load imbalance. We propose the first hybrid sparsity co-design framework tailored for heterogeneous Processing-in-Memory (PIM) architectures. By decoupling attention from feed-forward network computations, the framework integrates block-sparse attention, physical KV cache eviction, and adaptive expert routing. Leveraging a heterogeneous SRAM/HBM-PIM architecture, it further implements static expert mapping, dynamic sub-batch scheduling, and data path overlapping techniques. Experimental results demonstrate that the proposed approach achieves up to an 8.35× speedup over A100 GPUs and a 3.33× acceleration in FFN execution compared to PIMoE, all while preserving lossless inference accuracy.
📝 Abstract
Long-reasoning Mixture-of-Experts (MoE) models expose two coupled inference bottlenecks: growing KV caches shift the critical path toward attention, while sparse expert activation causes load imbalance and low hardware utilization. Although Processing-in-Memory (PIM) offers a promising way to mitigate data movement overhead, existing PIM-based accelerators typically optimize attention or FFNs in isolation. We propose SPIMOE, the first co-design framework that exploits hybrid sparsity for efficient MoE inference on heterogeneous PIM architectures. SPIMOE combines adaptive expert routing with block-sparse attention and physical KV-cache eviction, and disaggregates attention and FFNs across SRAM-PIM and HBM-PIM. Static expert mapping and dynamic sub-batch scheduling further balance channel loads and overlap the two paths. Evaluations show that SPIMOE achieves up to $8.35\times$ end-to-end speedup over an NVIDIA A100 GPU and $3.33\times$ speedup in MoE FFN execution over PIMoE, while preserving reasoning accuracy comparable to full-attention baselines.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Processing-in-Memory
KV cache
load imbalance
hardware utilization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Processing-in-Memory
Hybrid Sparsity
Hardware-Software Co-design
KV Cache Eviction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Rubing Yang
Rubing Yang
University of Pennsylvania
Deep learningMachine perception
C
Cenlin Duan
School of Integrated Circuits and Systems, Beihang University, Beijing, China
Y
Yingjie Qi
School of Computer Science, Beihang University, Beijing, China
X
Xiaolin He
School of Computer Science, Beihang University, Beijing, China
X
Xiao Ma
School of Computer Science, Beihang University, Beijing, China
Jianlei Yang
Jianlei Yang
Beihang University
Deep LearningComputer ArchitectureNueromorphic ComputingSpitronicsEDA/VLSI