A Machine Learning Approach Towards Runtime Optimisation of Matrix Multiplication

📅 2023-05-01
🏛️ IEEE International Parallel and Distributed Processing Symposium
📈 Citations: 2
Influential: 0
📄 PDF

career value

275K/year
🤖 AI Summary
This work addresses the challenge of dynamically determining the optimal thread count for General Matrix Multiplication (GEMM) on multi-core shared-memory systems. To this end, the authors propose the Architecture- and Data Structure-Aware Linear Algebra library (ADSALA), which integrates a runtime machine learning model into a BLAS library for the first time. By jointly considering task characteristics, hardware context, and conventional optimizations such as blocking, ADSALA dynamically selects the best thread configuration. Evaluated on dual-socket HPC nodes featuring Intel Cascade Lake and AMD Zen 3 architectures, the approach achieves a 25%–40% speedup over traditional BLAS implementations for GEMM workloads with memory footprints under 100 MB, significantly enhancing performance adaptability.

Technology Category

Application Category

📝 Abstract
The GEneral Matrix Multiplication (GEMM) is one of the essential algorithms in scientific computing. Single-thread GEMM implementations are well-optimised with techniques like blocking and autotuning. However, due to the complexity of modern multi-core shared memory systems, it is challenging to determine the number of threads that minimises the multi-thread GEMM runtime.We present a proof-of-concept approach to building an Architecture and Data-Structure Aware Linear Algebra (ADSALA) software library that uses machine learning to optimise the runtime performance of BLAS routines. More specifically, our method uses a machine learning model on-the-fly to automatically select the optimal number of threads for a given GEMM task based on the collected training data. Test results on two different HPC node architectures, one based on a two-socket Intel Cascade Lake and the other on a two-socket AMD Zen 3, revealed a 25 to 40 per cent speedup compared to traditional GEMM implementations in BLAS when using GEMM of memory usage within 100 MB.
Problem

Research questions and friction points this paper is trying to address.

GEMM
multi-threading
runtime optimisation
shared memory systems
optimal thread count
Innovation

Methods, ideas, or system contributions that make the work stand out.

machine learning
GEMM
runtime optimisation
thread selection
architecture-aware
🔎 Similar Papers
No similar papers found.