🤖 AI Summary
This study addresses the challenges of scheduling and accelerating data analytics tasks in heterogeneous superclouds by proposing ReDSEa, an automated compiler toolchain. Built upon LLVM, the method introduces novel performance models for recursive, iterative, and tiled computations, enabling end-to-end automation of task mapping, load balancing, and parallel execution. The system is deployed on a heterogeneous platform comprising Huawei Kunpeng 920 CPUs and Ascend 910 AI accelerators. Experimental evaluations demonstrate that, compared to a 48-core CPU baseline, the proposed toolchain achieves a 17× speedup for Cholesky (CH) decomposition and an 88× speedup for Singular Value Decomposition (SVD). These results establish ReDSEa as an efficient compilation optimization solution for scientific computing in heterogeneous cloud environments.
📝 Abstract
Heterogeneous Supercloud systems are transforming data analytics by enabling scalable and efficient task distribution across diverse resources. This paper presents computational models and performance estimation techniques tailored for accelerating Dense Cholesky (CH) Decomposition and Singular Value Decomposition (SVD)-based analytics on Heterogeneous Superclouds. Our ReDSEa tool-chain automates mapping, load balancing, scheduling, parallelism, and computation overlap. Implemented on a heterogeneous system with a Huawei Kunpeng 920 ARM CPU and an Ascend 910 AI accelerator, our LLVM compiler tool-chain employs novel performance models for recursive, iterative, and blocked computations, achieving up to 17x speedup for CH and 88x for SVD over fully optimized 48-core CPU implementations.