🤖 AI Summary
This work addresses the challenge of dynamically determining the optimal thread count for General Matrix Multiplication (GEMM) on multi-core shared-memory systems. To this end, the authors propose the Architecture- and Data Structure-Aware Linear Algebra library (ADSALA), which integrates a runtime machine learning model into a BLAS library for the first time. By jointly considering task characteristics, hardware context, and conventional optimizations such as blocking, ADSALA dynamically selects the best thread configuration. Evaluated on dual-socket HPC nodes featuring Intel Cascade Lake and AMD Zen 3 architectures, the approach achieves a 25%–40% speedup over traditional BLAS implementations for GEMM workloads with memory footprints under 100 MB, significantly enhancing performance adaptability.
📝 Abstract
The GEneral Matrix Multiplication (GEMM) is one of the essential algorithms in scientific computing. Single-thread GEMM implementations are well-optimised with techniques like blocking and autotuning. However, due to the complexity of modern multi-core shared memory systems, it is challenging to determine the number of threads that minimises the multi-thread GEMM runtime.We present a proof-of-concept approach to building an Architecture and Data-Structure Aware Linear Algebra (ADSALA) software library that uses machine learning to optimise the runtime performance of BLAS routines. More specifically, our method uses a machine learning model on-the-fly to automatically select the optimal number of threads for a given GEMM task based on the collected training data. Test results on two different HPC node architectures, one based on a two-socket Intel Cascade Lake and the other on a two-socket AMD Zen 3, revealed a 25 to 40 per cent speedup compared to traditional GEMM implementations in BLAS when using GEMM of memory usage within 100 MB.