Score
Design and implement compilers, DSLs, and code-generation pipelines that lower high-level sparse-tensor computations into efficient parallel code, including support for multiple sparse inputs and sparse outputs, representation-specific iteration and memory mappings, and composition of parallel loops. Build the analyses, lowering passes, and runtime support needed to produce and optimize code that matches or exceeds hand-optimized performance on target architectures.
Symmetric sparse tensor computations suffer from the difficulty of jointly optimizing symmetry and sparsity, while manual implementation is error-prone and combinatorially explosive. To address this, we present the first compiler that automatically generates functionally complete, symmetry-aware sparse tensor kernels. We introduce a novel taxonomy of tensor symmetries, unifying symmetry constraints with sparse iteration logic. We design a domain-specific intermediate representation (DSIR) and a symmetry-aware scheduler that integrate triangular loop clipping, transposition-equivalence class enumeration, and adaptive traversal across sparse formats. Evaluated on representative kernels—including SSYMV and 5D MTTKRP—our approach achieves speedups of 1.36×–30.4× over state-of-the-art asymmetric methods, demonstrating substantial performance gains and correctness guarantees through automated, symmetry-preserving code generation.
Existing sparse tensor compilers struggle to achieve safe and efficient parallelization in scenarios involving sparsity in the output or multiple sparse inputs. To address this challenge, this work introduces WingSpan, a sparse tensor language that, for the first time, supports unrestricted combinations of parallel loops and data structures. It also develops a concurrency dependency theory tailored for both sparse and structured tensors, ensuring correctness of parallel execution. By unifying these advances, the proposed framework delivers general-purpose parallel support for sparse tensor programs, matching or exceeding the performance of hand-optimized parallel implementations on key kernels such as sparse general matrix-matrix multiplication (SpGEMM).
Existing sparse and structured tensor computation frameworks suffer from rigid control-flow abstractions and fragmented structural support, hindering efficient exploitation of intrinsic properties such as sparsity, symmetry, and blocking. This paper introduces Finch—a novel domain-specific language that unifies arbitrary control flow (e.g., loops, conditionals, breaks) with diverse tensor structures (sparse, symmetric, blocked) via a joint control-flow–data-structure representation, enabling automatic structural specialization. Finch integrates structure-aware code generation, sparse tensor algebra compilation (e.g., SpMV, SpGEMM), and metadata-driven runtime execution. Evaluated on sparse matrix multiplication, image processing, and graph analytics, Finch achieves substantial performance gains over state-of-the-art frameworks. It significantly improves utilization of structural zeros, redundant values, and non-zero clusters—demonstrating superior efficiency in leveraging inherent tensor structure.
This work addresses efficient decomposition of high-dimensional sparse tensors—common in healthcare and cybersecurity—on modern parallel processors, overcoming restrictive assumptions about mode structure or sparsity distribution inherent in conventional compressed formats. We propose ALTO, an adaptive linearization tensor representation that is agnostic to both mode structure and sparsity distribution. Built upon ALTO, we design a parallel decomposition algorithm featuring low synchronization overhead and high data reuse, augmented by dynamic performance modeling and scheduling heuristics for automatic hardware adaptation. Leveraging cache- and memory-aware optimizations on Intel Xeon Scalable platforms, experiments demonstrate that ALTO achieves over 10× speedup versus the best structure-agnostic format and a 5.1× geometric mean speedup versus the best structure-aware format, while incurring only 25% of the latter’s storage overhead.
High-performance computing (HPC) stencil domain-specific language (DSL) compilers suffer from high development costs, poor infrastructure reuse, and low maintainability due to isolated, ad hoc designs. To address these challenges, this paper proposes MLIR-HPC, a dedicated extensible compiler framework for HPC built on the MLIR infrastructure. Our method introduces three key innovations: (1) a novel message-passing abstraction for distributed-memory systems that uniformly models communication semantics; (2) a distributed stencil intermediate representation (IR) supporting automated communication generation and cross-DSL optimization passes; and (3) seamless integration with three major DSL backends—Devito, PSyclone, and Open Earth Compiler—enabling shared compilation stack infrastructure. Evaluated across heterogeneous supercomputing architectures, the framework supports all three stencil DSLs using a unified core, achieving industrial-grade compilation efficiency and execution performance. Results demonstrate significantly enhanced sustainability, reusability, and evolutionary capability for HPC DSL compilers.
本文提出了一种新的代码生成策略,通过扩展TACO编译器的表示方法来处理稀疏张量收缩中的冲突数据布局问题,避免了显式转置,从而提高了计算效率。
Splyce通过双路径执行模型解决稀疏张量收缩中的矢量化难题,优化了内存延迟和指令级并行性,显著提升了性能。
Existing scientific computing codes are difficult to efficiently port to specialized architectures such as AMD AI Engine, often requiring extensive manual refactoring. This work proposes a tensor abstraction–based compilation approach that automatically elevates generic loops to tensor semantics by parsing lightweight OpenMP annotations, and constructs an end-to-end compilation pipeline to map computations onto the AI Engine execution model. The method significantly reduces programming complexity through minimal OpenMP directives and enables CPU–NPU cooperative scheduling. Experimental results on six scientific and AI kernel benchmarks show that the NPU achieves higher energy efficiency than a multi-core CPU at float32 precision; for two kernels, cooperative execution yields a 40% performance improvement and 15% energy reduction.
This work addresses the limited flexibility in distributed programming for large language model scaling and the inefficiency of existing tensor compilers in handling the complex memory hierarchies of heterogeneous clusters. To overcome these challenges, the authors propose a scalable block-level compiler featuring a novel three-tier hierarchical abstraction—Core, Device, and Task—that uniformly supports diverse parallelization strategies, automatically optimizes intra- and inter-node communication, and enables efficient code generation across both NVIDIA and AMD platforms. When integrated into vLLM, the compiler achieves 5%–30% end-to-end inference speedup and over 10% improvement in training model FLOPs utilization (MFU), translating to approximately 500,000 GPU hours saved per month. The system has been deployed in enterprise settings, delivering over 20% inference performance gains.
This work addresses the tight coupling between OpenMP semantics and fixed lowering strategies in existing compilers, which limits cross-hardware performance portability. Building upon the MLIR framework, we propose a modular parallel code generation approach that leverages domain-specific languages to declaratively specify lowering logic. By refactoring the lowering process into programmable components, our method decouples frontend semantics from backend targets, enabling explicit control over code outlining, data sharing, and runtime interfaces. Evaluations on the PolyBench/C-OMP benchmark suite demonstrate that this approach matches the performance of state-of-the-art compilers while introducing less than 0.7% code overhead. Furthermore, it reduces lowering code volume by 32% and 76% compared to Clang and GCC, respectively, and facilitates rapid adaptation to new runtimes.