Score
Designs and implements elementwise computational kernels expressed as tensorized operations using tensor programming (e.g., PyTorch); builds optimized per‑element tensor routines that avoid global assembly, minimize memory movement and compute overhead, and integrate with automatic differentiation.
This work addresses the challenge of deploying mainstream deep learning frameworks in resource-constrained environments, where their large size and lack of lightweight yet fully featured alternatives pose significant limitations. To this end, we propose and implement a lightweight tensor computation library built in Rust, leveraging its performance and memory safety guarantees to construct an efficient computational engine. The system exposes a PyTorch-like Python interface via PyO3, supporting essential features including n-dimensional tensors, dynamic computation graphs, reverse-mode automatic differentiation, neural network layers, and optimizers. The resulting installable package occupies only a few megabytes—orders of magnitude smaller than PyTorch or TensorFlow—while retaining the core capabilities necessary for research and development on CPU-based systems.
This work addresses the challenge of automatically generating high-performance GPU tiled kernels from high-level tensor algebra expressions, thereby alleviating the burden of manual optimization. The authors propose an end-to-end compilation framework that integrates layer-wise lowering, expression rewriting, automated schedule search, reduction fusion, and tiling optimizations. For the first time, this framework automatically discovers high-efficiency kernels—comparable to FlashAttention-3—from the mathematical specification of attention operators, while introducing a novel scheduler that preserves program structural regularity. Evaluated on GH200 and RTX 5090 GPUs, the generated kernels achieve up to 23% and 42% higher throughput, respectively, and match or surpass hand-optimized cuDNN kernels across multiple long-sequence configurations.
This work addresses the high barrier to entry in manually developing high-performance tensor kernel functions for AI accelerators, a process that traditionally demands deep expertise in tiling strategies, instruction selection, data layout optimization, and operator fusion. To alleviate this burden, the paper proposes an automated synthesis approach that eliminates the need for handcrafted rules by integrating program synthesis, SMT solving, and algebraic transformations of computation graphs. The method formally verifies semantic equivalence over unbounded tensors and systematically explores the space of tiling strategies and instruction/operator fusion under hardware constraints. This enables the automatic generation of kernels that are not only semantically correct but also highly efficient, significantly reducing memory traffic and improving execution performance while substantially lowering the development effort required.
To address the challenges of programming tensor accelerators and underutilization of hardware performance, this paper proposes an LLM-driven automated code optimization framework. Methodologically, it introduces a novel structured two-stage prompting paradigm—comprising planning and generation—integrated with a plug-and-play optimization menu and a hardware-closed-loop feedback mechanism, enabling cross-operator strategy reuse. The approach synergistically combines domain-specific knowledge modeling, compiler scheduling, tensor computation abstractions, and advanced LLM prompt engineering. Evaluated on GEMM, convolution, and fine-grained linear algebra workloads, the framework achieves up to 5.6× speedup over vendor libraries and 1.4× over expert-tuned code; strategy reuse improves sampling efficiency, yielding an additional 24% acceleration. The core contribution is the first LLM-compiler-hardware co-optimization paradigm specifically designed for tensor accelerators.
Existing sparse and structured tensor computation frameworks suffer from rigid control-flow abstractions and fragmented structural support, hindering efficient exploitation of intrinsic properties such as sparsity, symmetry, and blocking. This paper introduces Finch—a novel domain-specific language that unifies arbitrary control flow (e.g., loops, conditionals, breaks) with diverse tensor structures (sparse, symmetric, blocked) via a joint control-flow–data-structure representation, enabling automatic structural specialization. Finch integrates structure-aware code generation, sparse tensor algebra compilation (e.g., SpMV, SpGEMM), and metadata-driven runtime execution. Evaluated on sparse matrix multiplication, image processing, and graph analytics, Finch achieves substantial performance gains over state-of-the-art frameworks. It significantly improves utilization of structural zeros, redundant values, and non-zero clusters—demonstrating superior efficiency in leveraging inherent tensor structure.
This work addresses the challenge of efficiently fusing multiple reduction kernels in ML compilers, where joint search spaces grow prohibitively large. We propose Cleave, a compiler whose core innovation lies in symbolically decoupling algebraic transformations from operator scheduling. It introduces a novel superoptimization mechanism based on symbolic shapes to circumvent equivalence-checking overhead, and designs a Split operator supporting dynamic split counts to parallelize reduction dimensions. By integrating iterative tiling with horizontal fusion, Cleave establishes an automated kernel generation pipeline. Experimental results demonstrate up to 2.8× speedup over the strongest baseline and a 5.9× reduction in compilation time. Furthermore, on production workloads, Cleave outperforms handcrafted FlashInfer backends, achieving average speedups of 1.4× to 1.7×.
为解决GPU编程中性能优化与安全性问题,提出Exo-GPU语言,通过将并行性和同步性作为顺序代码注解处理,保证功能等效同时便于性能调优。
This study addresses the performance limitations of GPU kernels generated by existing compilers and the shortcomings of LLM-based optimization approaches, which often overlook model structure and lack end-to-end verification. We propose a multi-agent collaborative framework that treats compiled models as structured artifacts, focusing specifically on Triton sub-kernel optimization. This work introduces a novel schedule-aware agent search mechanism, integrated with vendor library call protection and a four-stage gated cascaded verification pipeline encompassing static analysis, correctness checking, and performance gating to ensure both safety and efficacy. Evaluated on the KernelBench benchmark, our method achieves average speedups of 1.40×, 1.15×, and 1.07× at the L1, L2, and L3 levels, respectively, compared to torch.compile, significantly enhancing inference efficiency.
论文解决了GPU内核在机器学习系统中确定性和数值再现性的问题,通过黑盒重建、编译器强制执行和静态验证等方法来控制张量核心的逐位行为。
本文通过基于等价饱和的EqiForge方法,解决了高效GPU实现张量程序时高、低级优化难以联合扩展的问题,实现了显著的性能提升。