Comparing CPU and GPU compute of PERMANOVA on MI300A

📅 2025-05-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates CPU–GPU computational performance disparities for the memory-intensive statistical method PERMANOVA on the AMD MI300A heterogeneous accelerator. Leveraging its unified high-bandwidth memory (HBM) architecture, we implement a brute-force permutation scheme that eliminates explicit data movement between CPU and GPU. Contrary to expectations, our GPU implementation significantly outperforms a highly optimized CPU version—incorporating cache locality optimizations and simultaneous multithreading (SMT) tuning. Key contributions include: (1) uncovering an unexpectedly strong acceleration effect of SMT on PERMANOVA, challenging the conventional wisdom that memory-bound algorithms must rely solely on cache optimization; and (2) empirically validating that the MI300A’s unified-memory-plus-heterogeneous-cores paradigm simplifies hardware adaptation for statistical computing, enabling efficient, low-overhead heterogeneous acceleration for high-dimensional statistical analysis in bioinformatics and related domains.

Technology Category

Machine Learning: Hardware-aware MLData Mining & Knowledge Management: Scalability, Parallel & Distributed SystemsHumans and AI: Other Foundations of Human Computation & AI

Application Category

Graph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationSystems and Infrastructure for Web, Mobile and WoT: Web performance, measurement, and characterization
📝 Abstract
Comparing the tradeoffs of CPU and GPU compute for memory-heavy algorithms is often challenging, due to the drastically different memory subsystems on host CPUs and discrete GPUs. The AMD MI300A is an exception, since it sports both CPU and GPU cores in a single package, all backed by the same type of HBM memory. In this paper we analyze the performance of Permutational Multivariate Analysis of Variance (PERMANOVA), a non-parametric method that tests whether two or more groups of objects are significantly different based on a categorical factor. This method is memory-bound and has been recently optimized for CPU cache locality. Our tests show that GPU cores on the MI300A prefer the brute force approach instead, significantly outperforming the CPU-based implementation. The significant benefit of Simultaneous Multithreading (SMT) was also a pleasant surprise.
Problem

Research questions and friction points this paper is trying to address.

Comparing CPU and GPU performance for PERMANOVA on MI300A
Analyzing memory-bound PERMANOVA with shared HBM memory
Evaluating brute-force GPU vs cache-optimized CPU approaches
Innovation

Methods, ideas, or system contributions that make the work stand out.

AMD MI300A integrates CPU and GPU with HBM memory
PERMANOVA optimized for GPU brute force approach
SMT significantly boosts performance unexpectedly
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.