sparse tensor compilation

Design and implement compilers, DSLs, and code-generation pipelines that lower high-level sparse-tensor computations into efficient parallel code, including support for multiple sparse inputs and sparse outputs, representation-specific iteration and memory mappings, and composition of parallel loops. Build the analyses, lowering passes, and runtime support needed to produce and optimize code that matches or exceeds hand-optimized performance on target architectures.

sparsetensorcompilation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

SySTeC: A Symmetric Sparse Tensor Compiler

Jun 13, 2024
RP
Radha Patel
🏛️ MIT

Symmetric sparse tensor computations suffer from the difficulty of jointly optimizing symmetry and sparsity, while manual implementation is error-prone and combinatorially explosive. To address this, we present the first compiler that automatically generates functionally complete, symmetry-aware sparse tensor kernels. We introduce a novel taxonomy of tensor symmetries, unifying symmetry constraints with sparse iteration logic. We design a domain-specific intermediate representation (DSIR) and a symmetry-aware scheduler that integrate triangular loop clipping, transposition-equivalence class enumeration, and adaptive traversal across sparse formats. Evaluated on representative kernels—including SSYMV and 5D MTTKRP—our approach achieves speedups of 1.36×–30.4× over state-of-the-art asymmetric methods, demonstrating substantial performance gains and correctness guarantees through automated, symmetry-preserving code generation.

Compiling symmetric sparse tensor kernels automaticallyGenerating optimized code for multiple tensor transpositionsHandling combinatorial symmetry and sparsity cases efficiently

Existing sparse tensor compilers struggle to achieve safe and efficient parallelization in scenarios involving sparsity in the output or multiple sparse inputs. To address this challenge, this work introduces WingSpan, a sparse tensor language that, for the first time, supports unrestricted combinations of parallel loops and data structures. It also develops a concurrency dependency theory tailored for both sparse and structured tensors, ensuring correctness of parallel execution. By unifying these advances, the proposed framework delivers general-purpose parallel support for sparse tensor programs, matching or exceeding the performance of hand-optimized parallel implementations on key kernels such as sparse general matrix-matrix multiplication (SpGEMM).

dependence analysisparallelismsparse tensors

Finch: Sparse and Structured Tensor Programming with Control Flow

Apr 25, 2024
WA
Willow Ahrens
🏛️ MIT | University of Washington

Existing sparse and structured tensor computation frameworks suffer from rigid control-flow abstractions and fragmented structural support, hindering efficient exploitation of intrinsic properties such as sparsity, symmetry, and blocking. This paper introduces Finch—a novel domain-specific language that unifies arbitrary control flow (e.g., loops, conditionals, breaks) with diverse tensor structures (sparse, symmetric, blocked) via a joint control-flow–data-structure representation, enabling automatic structural specialization. Finch integrates structure-aware code generation, sparse tensor algebra compilation (e.g., SpMV, SpGEMM), and metadata-driven runtime execution. Evaluated on sparse matrix multiplication, image processing, and graph analytics, Finch achieves substantial performance gains over state-of-the-art frameworks. It significantly improves utilization of structural zeros, redundant values, and non-zero clusters—demonstrating superior efficiency in leveraging inherent tensor structure.

Complex Data StructuresEfficient Tensor ProcessingOptimization of Computational Methods

Accelerating Sparse Tensor Decomposition Using Adaptive Linearized Representation

Mar 11, 2024
JL
Jan Laukemann
🏛️ Friedrich-Alexander-Universität Erlangen-Nüernberg | Intel Labs | University of Oregon | Laboratory for Physical Sciences

This work addresses efficient decomposition of high-dimensional sparse tensors—common in healthcare and cybersecurity—on modern parallel processors, overcoming restrictive assumptions about mode structure or sparsity distribution inherent in conventional compressed formats. We propose ALTO, an adaptive linearization tensor representation that is agnostic to both mode structure and sparsity distribution. Built upon ALTO, we design a parallel decomposition algorithm featuring low synchronization overhead and high data reuse, augmented by dynamic performance modeling and scheduling heuristics for automatic hardware adaptation. Leveraging cache- and memory-aware optimizations on Intel Xeon Scalable platforms, experiments demonstrate that ALTO achieves over 10× speedup versus the best structure-agnostic format and a 5.1× geometric mean speedup versus the best structure-aware format, while incurring only 25% of the latter’s storage overhead.

Efficient decomposition of high-dimensional sparse tensorsOvercoming irregular shapes and data distributions in sparse tensorsReducing memory footprint and synchronization overhead in tensor computations

A shared compilation stack for distributed-memory parallelism in stencil DSLs

Apr 02, 2024
GB
George Bisbas
🏛️ Imperial College London | The University of Edinburgh | Technische Universität Berlin | University of Cambridge

High-performance computing (HPC) stencil domain-specific language (DSL) compilers suffer from high development costs, poor infrastructure reuse, and low maintainability due to isolated, ad hoc designs. To address these challenges, this paper proposes MLIR-HPC, a dedicated extensible compiler framework for HPC built on the MLIR infrastructure. Our method introduces three key innovations: (1) a novel message-passing abstraction for distributed-memory systems that uniformly models communication semantics; (2) a distributed stencil intermediate representation (IR) supporting automated communication generation and cross-DSL optimization passes; and (3) seamless integration with three major DSL backends—Devito, PSyclone, and Open Earth Compiler—enabling shared compilation stack infrastructure. Evaluated across heterogeneous supercomputing architectures, the framework supports all three stencil DSLs using a unified core, achieving industrial-grade compilation efficiency and execution performance. Results demonstrate significantly enhanced sustainability, reusability, and evolutionary capability for HPC DSL compilers.

Develop shared compiler infrastructure for distributed-memory parallelism in stencil DSLs.Enable high-performance executables across multiple HPC stencil-DSL compilers.Reduce development and maintenance costs of DSL compilers in HPC.

Latest Papers

What's happening recently
View more

Existing scientific computing codes are difficult to efficiently port to specialized architectures such as AMD AI Engine, often requiring extensive manual refactoring. This work proposes a tensor abstraction–based compilation approach that automatically elevates generic loops to tensor semantics by parsing lightweight OpenMP annotations, and constructs an end-to-end compilation pipeline to map computations onto the AI Engine execution model. The method significantly reduces programming complexity through minimal OpenMP directives and enables CPU–NPU cooperative scheduling. Experimental results on six scientific and AI kernel benchmarks show that the NPU achieves higher energy efficiency than a multi-core CPU at float32 precision; for two kernels, cooperative execution yields a 40% performance improvement and 15% energy reduction.

AI Enginescode portinghardware acceleration

This work addresses the limited flexibility in distributed programming for large language model scaling and the inefficiency of existing tensor compilers in handling the complex memory hierarchies of heterogeneous clusters. To overcome these challenges, the authors propose a scalable block-level compiler featuring a novel three-tier hierarchical abstraction—Core, Device, and Task—that uniformly supports diverse parallelization strategies, automatically optimizes intra- and inter-node communication, and enables efficient code generation across both NVIDIA and AMD platforms. When integrated into vLLM, the compiler achieves 5%–30% end-to-end inference speedup and over 10% improvement in training model FLOPs utilization (MFU), translating to approximately 500,000 GPU hours saved per month. The system has been deployed in enterprise settings, delivering over 20% inference performance gains.

distributed programminglarge language modelsmemory hierarchy

This work addresses the tight coupling between OpenMP semantics and fixed lowering strategies in existing compilers, which limits cross-hardware performance portability. Building upon the MLIR framework, we propose a modular parallel code generation approach that leverages domain-specific languages to declaratively specify lowering logic. By refactoring the lowering process into programmable components, our method decouples frontend semantics from backend targets, enabling explicit control over code outlining, data sharing, and runtime interfaces. Evaluations on the PolyBench/C-OMP benchmark suite demonstrate that this approach matches the performance of state-of-the-art compilers while introducing less than 0.7% code overhead. Furthermore, it reduces lowering code volume by 32% and 76% compared to Clang and GCC, respectively, and facilitates rapid adaptation to new runtimes.

compiler loweringhardware diversityOpenMP

Hot Scholars

HS

Heng Shi

Assistant Research Fellow of Tsinghua University; Visiting scholar at Technion
Cooperative guidanceAerospaceIntelligent control
JH

Junhui Hou

Department of Computer Science, City University of Hong Kong
Neural Spatial Computing
JC

Jianfei Chen

Associate Professor, Tsinghua University
Machine Learning
FK

Fredrik Kjolstad

Assistant Professor, Stanford University
CompilersProgramming LanguagesSparse ComputationPerformance Engineering
LH

Lin Huang

Stanford University, The Chinese University of Hong Kong
computational genomicsfault-tolerant computingdesign automationand multi-core architecture