🤖 AI Summary
This study systematically evaluates CPU–GPU computational performance disparities for the memory-intensive statistical method PERMANOVA on the AMD MI300A heterogeneous accelerator. Leveraging its unified high-bandwidth memory (HBM) architecture, we implement a brute-force permutation scheme that eliminates explicit data movement between CPU and GPU. Contrary to expectations, our GPU implementation significantly outperforms a highly optimized CPU version—incorporating cache locality optimizations and simultaneous multithreading (SMT) tuning. Key contributions include: (1) uncovering an unexpectedly strong acceleration effect of SMT on PERMANOVA, challenging the conventional wisdom that memory-bound algorithms must rely solely on cache optimization; and (2) empirically validating that the MI300A’s unified-memory-plus-heterogeneous-cores paradigm simplifies hardware adaptation for statistical computing, enabling efficient, low-overhead heterogeneous acceleration for high-dimensional statistical analysis in bioinformatics and related domains.
📝 Abstract
Comparing the tradeoffs of CPU and GPU compute for memory-heavy algorithms is often challenging, due to the drastically different memory subsystems on host CPUs and discrete GPUs. The AMD MI300A is an exception, since it sports both CPU and GPU cores in a single package, all backed by the same type of HBM memory. In this paper we analyze the performance of Permutational Multivariate Analysis of Variance (PERMANOVA), a non-parametric method that tests whether two or more groups of objects are significantly different based on a categorical factor. This method is memory-bound and has been recently optimized for CPU cache locality. Our tests show that GPU cores on the MI300A prefer the brute force approach instead, significantly outperforming the CPU-based implementation. The significant benefit of Simultaneous Multithreading (SMT) was also a pleasant surprise.