loop tiling

Designs, implements, or analyzes loop transformation passes or manual code changes that partition loop iteration spaces into tiles/blocks and that choose tile sizes, loop ordering, and scheduling to improve data locality, cache reuse, and exposure of parallelism and vectorization. Works with related transformations such as fusion, blocking, and tile-level culling, and evaluates trade-offs between memory footprint, redundant work, and runtime performance when integrating tiling with vectorization and scheduling.

looptiling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$210K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Unified Framework for Automated Code Transformation and Pragma Insertion

May 05, 2024
SP
Stéphane Pouget
🏛️ University of California, Los Angeles | Colorado State University

In high-level synthesis (HLS), jointly optimizing code transformations, pragma insertion, and cache-blocking size selection is challenging due to tight coupling, a vast decision space, and difficulty in guaranteeing semantic correctness. Method: This paper proposes the first unified modeling framework that jointly encodes all three aspects as a single, isomorphic optimization problem—supporting “zero-transformation” decisions—and leverages HLS compiler–driven constraint derivation coupled with nonlinear programming (NLP) to automatically and correctly optimize regular loop nests. Contribution/Results: It introduces the first paradigm for co-optimizing transformations, pragmas, and blocking sizes, with built-in semantic equivalence preservation. Evaluated on multiple benchmark kernels, the approach significantly improves quality-of-results (QoR), accurately identifies cases requiring or forbidding transformations, and generates high-performance, formally verifiable optimized code.

AutomationCode ModificationSimplification

This work presents the first systematic empirical study of code generation bugs in the Tile programming framework, which are often silent correctness or performance defects arising from tight coupling among input shapes, data types, and backend targets during multi-stage compilation—issues largely undetectable by existing testing tools. Through an in-depth analysis of 301 real-world bugs distilled from 401 GitHub reports, the study constructs the first comprehensive taxonomy of Tile-related defects, elucidating their root causes, triggering patterns, symptomatic manifestations, and effective repair strategies. The findings establish a foundational understanding that informs debugging, testing, and automated bug detection for Tile-based compiler infrastructures, offering both theoretical insights and practical guidance for improving their reliability and robustness.

code generation bugscompiler testingGPU kernels

LEGO: Layout Expression for Generating One-to-one Mapping

May 12, 2025
AM
Amir Mohammad Tavakkoli
🏛️ University of Utah | University of Copenhagen

To address the optimization limitations imposed by tight coupling between data layout and computation on GPUs, this paper proposes a layout-agnostic computational abstraction paradigm: computations are first expressed in a layout-decoupled form, and hierarchical parallel index expressions are then automatically derived via layout specifications. The core contribution is the first end-to-end “layout → index expression” mapping mechanism, enabling layout-driven code generation and cross-compiler optimization exploration. We design a custom layout specification language and integrate it with MLIR, Triton, and CUDA templates to build an index derivation engine. Experimental evaluation demonstrates that the generated code achieves performance on par with hand-optimized Triton kernels. Furthermore, the approach is validated for generality and efficiency across both MLIR-based and CUDA-based compilation ecosystems.

Exploring data layouts with compiler integrationGenerating complex indexing for hierarchical GPU codeOptimizing data movement via layout-independent computation

LOOPer: A Learned Automatic Code Optimizer For Polyhedral Compilers

Mar 18, 2024
MM
Massinissa Merouani
🏛️ New York University | École Nationale Supérieure d’Informatique | Meta AI

Existing polyhedral compilers suffer from limited support for deep affine transformations and rely on oversimplified program assumptions—such as rectangular iteration domains and single-level loop nests—hindering automatic selection of high-benefit schedules and impairing generality. This paper presents the first deep learning–driven clustering auto-scheduler designed for large-scale affine transformation spaces and complex program structures, including non-rectangular iteration domains and multi-level nested loops. Our approach integrates deep learning–based cost modeling, clustering-aware dependence analysis, multi-stage transformation sequence search, and iteration domain normalization with feature encoding. Evaluated on the PolyBench benchmark suite, our scheduler achieves geometric mean speedups of 1.84× over Tiramisu and 1.42× over Pluto. It significantly enhances optimization capability and practical applicability for complex, real-world programs.

Scaling optimization to non-rectangular and multi-loop programsSelecting optimal polyhedral transformations for speedupsSupporting complex affine transformations in compilers

TileLang: A Composable Tiled Programming Model for AI Systems

Apr 24, 2025
LW
Lei Wang
🏛️ Peking University | Imperial College London | Microsoft Research

AI kernel development faces challenges including complex hardware adaptation and insufficient expressiveness and usability of domain-specific compilers. This paper proposes a composable tiling programming model that—uniquely—decouples dataflow from the scheduling space (thread mapping, memory layout, tensorization, and pipelining). By introducing a unified block-thread abstraction, lightweight scheduling primitive annotations, and dataflow-driven, hardware-aware compilation, our approach jointly optimizes developer productivity and kernel performance. The method bridges the gap between expressive power and engineering practicality in domain-specific compilers. Evaluated on mainstream accelerators—including GPUs and ASICs—our generated AI compute kernels achieve state-of-the-art performance across key workloads, while reducing development cycles significantly. The framework delivers both flexibility in algorithmic expression and high execution efficiency.

Achieve state-of-the-art performance with easeDecouple scheduling from dataflow for flexibilitySimplify writing high-performance AI kernels

Latest Papers

What's happening recently
View more

This study addresses the limited transferability of optimization knowledge in dataflow architectures caused by reliance on vendor libraries. To this end, it proposes Loom, a symbolic compiler that formulates SPMD compilation as hardware-explicit static optimization while maintaining parameters in symbolic form. By jointly optimizing schedules and parameters through CP-SAT constraint solving, legality derivation, and spatial mapping enumeration, Loom establishes the first tuning-free symbolic compilation framework, enabling cross-architecture retargetable and interpretable compiler optimizations. Evaluated on the Tenstorrent architecture, Loom matches or surpasses vendor library performance on workloads such as GEMM without requiring shape-by-shape profiling.

auto-tuningcompiler schedulingdataflow architectures

This work addresses the performance bottlenecks in AI and high-performance computing caused by data movement in loop programs by proposing a domain-specific language (DSL)-based approach for automated locality analysis. The method formalizes affine loop nests as polyhedral sets and maps, enabling, for the first time, fully symbolic derivation of reuse distance and data movement complexity without relying on traditional techniques such as stack simulation or recursive working-set models. Implemented in Rust, the system integrates Barvinok counting, the polyhedral model, and affine transformations, and provides both a command-line tool and an interactive web platform. It enables precise locality analysis for representative operators including tensor contractions, einsum expressions, and stencil computations.

AI kernelsdata localitydata movement

This work addresses the challenge of energy estimation for nested-loop programs on parallel processor arrays, where traditional simulation-based approaches suffer from poor scalability. To overcome this limitation, the paper proposes a symbolic polyhedral energy modeling method that, for the first time, applies symbolic polyhedral analysis to energy estimation of nested loops. By integrating loop transformation theory with array architecture modeling, the approach explicitly captures the impact of mapping and scheduling decisions on energy consumption. Experimental results demonstrate that the method achieves high-accuracy energy predictions across multiple benchmarks, with computational overhead independent of problem size, thereby significantly enhancing the scalability of design space exploration.

energy analysismapping and schedulingnested loop programs

This study addresses the absence of systematic comparisons between tile-based programming models such as Triton and cuTile by constructing a controlled benchmark for standardized evaluation on NVIDIA B200 GPUs. Methodologically, we design a unified test suite comprising 45 operators and integrate PyTorch reference implementations, auto-tuning, Roofline modeling, and GPU profiling techniques for comprehensive bottleneck diagnosis. Furthermore, this work presents the first quantitative comparison of token efficiency for LLM-generated code. Our findings reveal that cuTile excels in Tensor Core-intensive kernels, whereas Triton demonstrates superior performance and greater token efficiency in irregular and memory-bandwidth-bound scenarios.

benchmarkbottleneck diagnosiskernel development

Hot Scholars

LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI
YL

Yansheng Li

Professor, Wuhan University
Deep LearningKnowledge GraphRemote Sensing Big Data Mining
GY

Guangwen Yang

Professor of Computer Science and Technology, Tsinghua University
CR

Caleb Robinson

Microsoft AI for Good
computational sustainabilitydeep learninghuman migration