CELLO: Co-designing Schedule and Hybrid Implicit/Explicit Buffer for Complex Tensor Reuse

📅 2023-03-20
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
Irregular tensor computations in high-performance computing (HPC)—such as those in conjugate gradient methods—suffer from low data reuse due to irregular operator shapes and complex dependency-directed acyclic graphs (DAGs), hindering co-optimization between conventional schedulers and hardware. To address this, we propose a software-hardware co-design acceleration architecture. Our key contributions are: (1) CHORD, a novel hybrid on-chip buffer mechanism supporting both implicit and explicit data management; (2) SCORE, a DAG-aware scheduler that enables fine-grained reuse exploitation across tensor operations in a DAG; and (3) tight integration of tensor algebra compilation optimizations with a customized microarchitecture. Evaluated on representative HPC workloads, our architecture achieves a 4.0× geometric mean speedup and 4.0× energy efficiency improvement over state-of-the-art accelerators, significantly overcoming the data reuse bottleneck in irregular tensor computations.
📝 Abstract
Tensor algebra accelerators have been gaining popularity for running high-performance computing (HPC) workloads. Identifying optimal schedules for individual tensor operations and designing hardware to run these schedules is an active area of research. Unfortunately, operators in HPC workloads such as Conjugate Gradient often have operators with skewed shapes, fundamentally limiting the reuse any schedule can leverage. Moreover, the operators form a complex DAG of dependencies, making it challenging to apply simple fusion/pipelining techniques to extract inter-operation reuse. To address these challenges, this work proposes an accelerator CELLO. CELLO uses a novel on-chip buffer mechanism called CHORD co-designed with a novel scheduler called SCORE, which together enables identifying and exploiting reuse over complex DAGs of tensor operations. CELLO provides 4x geomean speedup and 4x energy efficiency over state-of-the-art accelerators across HPC workloads.
Problem

Research questions and friction points this paper is trying to address.

Optimizing schedules for tensor operations in HPC workloads.
Addressing skewed tensor shapes limiting reuse in schedules.
Exploiting reuse in complex DAGs of tensor operations.
Innovation

Methods, ideas, or system contributions that make the work stand out.

CHORD: novel on-chip buffer mechanism
SCORE: co-designed scheduler for reuse
Exploits reuse in complex DAGs
🔎 Similar Papers
No similar papers found.
Georgia Tech | NVIDIA Research | Sandia National Laboratories
Raveesh Garg
Raveesh Garg
IBM
Computer ArchitectureHardware AcceleratorsProgrammable Spatial Architectures
M
Michael Pellauer
NVIDIA Research
S
S. Rajamanickam
Sandia National Laboratories
T
T. Krishna
Georgia Tech