tensorized elementwise kernels

Designs and implements elementwise computational kernels expressed as tensorized operations using tensor programming (e.g., PyTorch); builds optimized per‑element tensor routines that avoid global assembly, minimize memory movement and compute overhead, and integrate with automatic differentiation.

tensorizedelementwisekernels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.32
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of deploying mainstream deep learning frameworks in resource-constrained environments, where their large size and lack of lightweight yet fully featured alternatives pose significant limitations. To this end, we propose and implement a lightweight tensor computation library built in Rust, leveraging its performance and memory safety guarantees to construct an efficient computational engine. The system exposes a PyTorch-like Python interface via PyO3, supporting essential features including n-dimensional tensors, dynamic computation graphs, reverse-mode automatic differentiation, neural network layers, and optimizers. The resulting installable package occupies only a few megabytes—orders of magnitude smaller than PyTorch or TensorFlow—while retaining the core capabilities necessary for research and development on CPU-based systems.

CPU-based developmentdeep learning frameworkinstall footprint

This work addresses the challenge of automatically generating high-performance GPU tiled kernels from high-level tensor algebra expressions, thereby alleviating the burden of manual optimization. The authors propose an end-to-end compilation framework that integrates layer-wise lowering, expression rewriting, automated schedule search, reduction fusion, and tiling optimizations. For the first time, this framework automatically discovers high-efficiency kernels—comparable to FlashAttention-3—from the mathematical specification of attention operators, while introducing a novel scheduler that preserves program structural regularity. Evaluated on GH200 and RTX 5090 GPUs, the generated kernels achieve up to 23% and 42% higher throughput, respectively, and match or surpass hand-optimized cuDNN kernels across multiple long-sequence configurations.

auto-schedulingGPU kernel optimizationhigh-level optimization

This work addresses the high barrier to entry in manually developing high-performance tensor kernel functions for AI accelerators, a process that traditionally demands deep expertise in tiling strategies, instruction selection, data layout optimization, and operator fusion. To alleviate this burden, the paper proposes an automated synthesis approach that eliminates the need for handcrafted rules by integrating program synthesis, SMT solving, and algebraic transformations of computation graphs. The method formally verifies semantic equivalence over unbounded tensors and systematically explores the space of tiling strategies and instruction/operator fusion under hardware constraints. This enables the automatic generation of kernels that are not only semantically correct but also highly efficient, significantly reducing memory traffic and improving execution performance while substantially lowering the development effort required.

AI acceleratorshigh performance kernelsoperator fusion

Autocomp: LLM-Driven Code Optimization for Tensor Accelerators

May 24, 2025
CH
Charles Hong
🏛️ UC Berkeley

To address the challenges of programming tensor accelerators and underutilization of hardware performance, this paper proposes an LLM-driven automated code optimization framework. Methodologically, it introduces a novel structured two-stage prompting paradigm—comprising planning and generation—integrated with a plug-and-play optimization menu and a hardware-closed-loop feedback mechanism, enabling cross-operator strategy reuse. The approach synergistically combines domain-specific knowledge modeling, compiler scheduling, tensor computation abstractions, and advanced LLM prompt engineering. Evaluated on GEMM, convolution, and fine-grained linear algebra workloads, the framework achieves up to 5.6× speedup over vendor libraries and 1.4× over expert-tuned code; strategy reuse improves sampling efficiency, yielding an additional 24% acceleration. The core contribution is the first LLM-compiler-hardware co-optimization paradigm specifically designed for tensor accelerators.

Challenges in low-resource code generation for acceleratorsImproving performance via automated LLM-driven searchOptimizing tensor accelerator code using LLMs

Finch: Sparse and Structured Tensor Programming with Control Flow

Apr 25, 2024
WA
Willow Ahrens
🏛️ MIT | University of Washington

Existing sparse and structured tensor computation frameworks suffer from rigid control-flow abstractions and fragmented structural support, hindering efficient exploitation of intrinsic properties such as sparsity, symmetry, and blocking. This paper introduces Finch—a novel domain-specific language that unifies arbitrary control flow (e.g., loops, conditionals, breaks) with diverse tensor structures (sparse, symmetric, blocked) via a joint control-flow–data-structure representation, enabling automatic structural specialization. Finch integrates structure-aware code generation, sparse tensor algebra compilation (e.g., SpMV, SpGEMM), and metadata-driven runtime execution. Evaluated on sparse matrix multiplication, image processing, and graph analytics, Finch achieves substantial performance gains over state-of-the-art frameworks. It significantly improves utilization of structural zeros, redundant values, and non-zero clusters—demonstrating superior efficiency in leveraging inherent tensor structure.

Complex Data StructuresEfficient Tensor ProcessingOptimization of Computational Methods

Latest Papers

What's happening recently
View more

This work addresses the challenge of efficiently fusing multiple reduction kernels in ML compilers, where joint search spaces grow prohibitively large. We propose Cleave, a compiler whose core innovation lies in symbolically decoupling algebraic transformations from operator scheduling. It introduces a novel superoptimization mechanism based on symbolic shapes to circumvent equivalence-checking overhead, and designs a Split operator supporting dynamic split counts to parallelize reduction dimensions. By integrating iterative tiling with horizontal fusion, Cleave establishes an automated kernel generation pipeline. Experimental results demonstrate up to 2.8× speedup over the strongest baseline and a 5.9× reduction in compilation time. Furthermore, on production workloads, Cleave outperforms handcrafted FlashInfer backends, achieving average speedups of 1.4× to 1.7×.

Kernel GenerationML CompilerOperator Fusion

This study addresses the performance limitations of GPU kernels generated by existing compilers and the shortcomings of LLM-based optimization approaches, which often overlook model structure and lack end-to-end verification. We propose a multi-agent collaborative framework that treats compiled models as structured artifacts, focusing specifically on Triton sub-kernel optimization. This work introduces a novel schedule-aware agent search mechanism, integrated with vendor library call protection and a four-stage gated cascaded verification pipeline encompassing static analysis, correctness checking, and performance gating to ensure both safety and efficacy. Evaluated on the KernelBench benchmark, our method achieves average speedups of 1.40×, 1.15×, and 1.07× at the L1, L2, and L3 levels, respectively, compared to torch.compile, significantly enhancing inference efficiency.

deep learning compilerend-to-end verificationGPU kernel optimization

Hot Scholars

JT

Jacek Tabor

Profesor informatyki, Uniwersytet Jagielloński
mathematicscomputer science
MI

Måns I. Andersson

Virginia Tech | KTH Royal Institute of Technology, PhD
PDEHPCScientific ComputingParallel Computing
HT

Hao-Ting Pai

National Pingtung University
AI's BiasDisparityand InterpretabilityMisdiagnosis
JS

Justin Solomon

MIT
Computer graphicsgeometry processingmachine learning
JN

Jannatun Noor

Associate Professor, CSE, and Director, BSDS, UIU
Cloud ComputingICTDICT4DCSCW