🤖 AI Summary
This study addresses the high latency of Mixture-of-Experts (MoE) inference on CPUs, which is constrained by memory bandwidth, and the poor portability caused by the intrusive nature of existing acceleration schemes. To this end, we propose a non-intrusive MoE CPU inference library. Methodologically, micro-kernels are optimized via the Intel AMX instruction set to transform matrix-vector operations into matrix-matrix computations. Additionally, a lightweight token transformation strategy that preserves standard weight layouts is designed, combined with NUMA-aware task partitioning and a framework-agnostic page interleaving mechanism. This approach achieves plug-and-play compatibility, delivering up to 3.37× speedup for FFN kernels and up to 2.09× end-to-end inference acceleration.
📝 Abstract
Mixture-of-Experts (MoE) inference is increasingly deployed in local and on-premise environments, where expert parameters often exceed GPU memory capacity. In latency-sensitive, low-concurrency settings, repeatedly staging routed-expert weights from CPU memory to the GPU can be prohibitive, leaving routed-expert feed-forward networks (FFNs) on the critical path of multi-socket CPUs. Existing CPU accelerations often rely on intrusive, hardware- or topology-specific requirements, such as AMX-specific weight layouts or manual NUMA-aware placement. These requirements reduce portability and complicate integration with standard CPU-GPU offloading pipelines. We present HiNa-MoE, a high-performance, non-intrusive operator library for MoE inference on CPUs with Intel AMX. HiNa-MoE (1) exploits AMX with an optimized micro-kernel that keeps expert weights in standard layouts and instead fuses lightweight layout transforms into token gathering and stores; (2) applies NUMA-aware task partitioning under a simple page-interleaved policy without modifying the framework allocator; and (3) converts decode-phase memory matrix-vector operations into small matrix-matrix execution to utilize AMX. Across multiple MoE models, HiNa-MoE achieves up to 3.37x speedup for FFN kernels and up to 2.09x end-to-end inference speedup over state-of-the-art baselines, while remaining plug-and-play with existing frameworks and deployment workflows.