🤖 AI Summary
This study addresses the hardware utilization bottlenecks in long-reasoning Mixture-of-Experts (MoE) models caused by KV cache bloat and expert load imbalance. We propose the first hybrid sparsity co-design framework tailored for heterogeneous Processing-in-Memory (PIM) architectures. By decoupling attention from feed-forward network computations, the framework integrates block-sparse attention, physical KV cache eviction, and adaptive expert routing. Leveraging a heterogeneous SRAM/HBM-PIM architecture, it further implements static expert mapping, dynamic sub-batch scheduling, and data path overlapping techniques. Experimental results demonstrate that the proposed approach achieves up to an 8.35× speedup over A100 GPUs and a 3.33× acceleration in FFN execution compared to PIMoE, all while preserving lossless inference accuracy.
📝 Abstract
Long-reasoning Mixture-of-Experts (MoE) models expose two coupled inference bottlenecks: growing KV caches shift the critical path toward attention, while sparse expert activation causes load imbalance and low hardware utilization. Although Processing-in-Memory (PIM) offers a promising way to mitigate data movement overhead, existing PIM-based accelerators typically optimize attention or FFNs in isolation. We propose SPIMOE, the first co-design framework that exploits hybrid sparsity for efficient MoE inference on heterogeneous PIM architectures. SPIMOE combines adaptive expert routing with block-sparse attention and physical KV-cache eviction, and disaggregates attention and FFNs across SRAM-PIM and HBM-PIM. Static expert mapping and dynamic sub-batch scheduling further balance channel loads and overlap the two paths. Evaluations show that SPIMOE achieves up to $8.35\times$ end-to-end speedup over an NVIDIA A100 GPU and $3.33\times$ speedup in MoE FFN execution over PIMoE, while preserving reasoning accuracy comparable to full-attention baselines.