apply register blocking

Design and implement loop-tiling and scheduling transformations that partition computations into small blocks sized to fit in processor registers (register blocking), so inner-loop operations work on register-resident tiles and minimize loads/stores to higher-level memory. Build the code, data layouts, and analyses that choose block sizes and instruction sequences to stream data through on‑chip registers, reduce peak memory footprint, and avoid materializing full intermediate matrices.

applyregisterblocking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Unified Framework for Automated Code Transformation and Pragma Insertion

May 05, 2024
SP
Stéphane Pouget
🏛️ University of California, Los Angeles | Colorado State University

In high-level synthesis (HLS), jointly optimizing code transformations, pragma insertion, and cache-blocking size selection is challenging due to tight coupling, a vast decision space, and difficulty in guaranteeing semantic correctness. Method: This paper proposes the first unified modeling framework that jointly encodes all three aspects as a single, isomorphic optimization problem—supporting “zero-transformation” decisions—and leverages HLS compiler–driven constraint derivation coupled with nonlinear programming (NLP) to automatically and correctly optimize regular loop nests. Contribution/Results: It introduces the first paradigm for co-optimizing transformations, pragmas, and blocking sizes, with built-in semantic equivalence preservation. Evaluated on multiple benchmark kernels, the approach significantly improves quality-of-results (QoR), accurately identifies cases requiring or forbidding transformations, and generates high-performance, formally verifiable optimized code.

AutomationCode ModificationSimplification

This study addresses the limited transferability of optimization knowledge in dataflow architectures caused by reliance on vendor libraries. To this end, it proposes Loom, a symbolic compiler that formulates SPMD compilation as hardware-explicit static optimization while maintaining parameters in symbolic form. By jointly optimizing schedules and parameters through CP-SAT constraint solving, legality derivation, and spatial mapping enumeration, Loom establishes the first tuning-free symbolic compilation framework, enabling cross-architecture retargetable and interpretable compiler optimizations. Evaluated on the Tenstorrent architecture, Loom matches or surpasses vendor library performance on workloads such as GEMM without requiring shape-by-shape profiling.

auto-tuningcompiler schedulingdataflow architectures

LOOPer: A Learned Automatic Code Optimizer For Polyhedral Compilers

Mar 18, 2024
MM
Massinissa Merouani
🏛️ New York University | École Nationale Supérieure d’Informatique | Meta AI

Existing polyhedral compilers suffer from limited support for deep affine transformations and rely on oversimplified program assumptions—such as rectangular iteration domains and single-level loop nests—hindering automatic selection of high-benefit schedules and impairing generality. This paper presents the first deep learning–driven clustering auto-scheduler designed for large-scale affine transformation spaces and complex program structures, including non-rectangular iteration domains and multi-level nested loops. Our approach integrates deep learning–based cost modeling, clustering-aware dependence analysis, multi-stage transformation sequence search, and iteration domain normalization with feature encoding. Evaluated on the PolyBench benchmark suite, our scheduler achieves geometric mean speedups of 1.84× over Tiramisu and 1.42× over Pluto. It significantly enhances optimization capability and practical applicability for complex, real-world programs.

Scaling optimization to non-rectangular and multi-loop programsSelecting optimal polyhedral transformations for speedupsSupporting complex affine transformations in compilers

TileLang: A Composable Tiled Programming Model for AI Systems

Apr 24, 2025
LW
Lei Wang
🏛️ Peking University | Imperial College London | Microsoft Research

AI kernel development faces challenges including complex hardware adaptation and insufficient expressiveness and usability of domain-specific compilers. This paper proposes a composable tiling programming model that—uniquely—decouples dataflow from the scheduling space (thread mapping, memory layout, tensorization, and pipelining). By introducing a unified block-thread abstraction, lightweight scheduling primitive annotations, and dataflow-driven, hardware-aware compilation, our approach jointly optimizes developer productivity and kernel performance. The method bridges the gap between expressive power and engineering practicality in domain-specific compilers. Evaluated on mainstream accelerators—including GPUs and ASICs—our generated AI compute kernels achieve state-of-the-art performance across key workloads, while reducing development cycles significantly. The framework delivers both flexibility in algorithmic expression and high execution efficiency.

Achieve state-of-the-art performance with easeDecouple scheduling from dataflow for flexibilitySimplify writing high-performance AI kernels

Automatic Hardware Pragma Insertion in High-Level Synthesis: A Non-Linear Programming Approach

Apr 01, 2024
SP
Stéphane Pouget
🏛️ University of California, Los Angeles | Colorado State University

Manual pragma configuration in high-level synthesis (HLS) suffers from low efficiency and an exponentially large search space. Method: This paper proposes the first nonlinear programming (NLP)-based automated pragma insertion framework, jointly optimizing loop-level pragmas—including pipelining, function unit replication, and data caching. It innovatively models discrete pragma configurations as continuous, differentiable variables and constructs analytical performance/resource models with theoretical lower-bound guarantees, solved globally via NLP. Integrated with pragma semantic analysis and the Merlin compiler, and augmented by design-space pruning, the framework explores billion-scale configurations within seconds to minutes. Contribution/Results: Experimental evaluation shows kernel performance approaching hand-tuned implementations, resource estimation error <8%, and latency lower-bound error ≤12%.

Automatic insertion of HLS pragmasNon-linear programming for hardware synthesisOptimizing pipelining and data caching

Latest Papers

What's happening recently
View more

This work addresses the lack of efficient and reusable compilation infrastructure between high-level AI frameworks and hardware accelerators, which hinders automatic generation of high-performance code. Building upon MLIR, the authors propose a modular compiler featuring a lightweight affine analysis pipeline that integrates loop transformations, multi-level tiling, operator and attention-layer fusion, on-chip memory management, and mapping to specialized compute units. Combined with an analytical cost model and heuristic strategies, the system enables fully automated optimization from PyTorch/JAX down to hardware primitives. Evaluated on NVIDIA GPUs, the JIT-compiled code matches or exceeds the performance of Torch Inductor and XLA, with generated matrix multiplication and convolution kernels achieving parity with vendor-optimized libraries or hand-tuned kernels.

AI chipsAI programming frameworksautomatic code generation

This work proposes a method to automatically derive efficient tiling and prefetching schedules from hardware cache hierarchies, eliminating reliance on empirical tuning. Building upon the Mathematics of Arrays framework, it introduces a machine shape—characterized by cache capacities, bandwidths, and occupancy sequences—to model multilevel caches, and designs novel operators that map this machine shape to tiling strategies, reducing unknown parameters to a small set of level-specific occupancies. The study addresses the fundamental open question of which parameters can be directly inferred from hardware specifications. Experiments on three real machines successfully reproduce expert-tuned tile sizes, and demonstrate that certain throughput-related parameters can indeed be derived from datasheets; however, cross-architecture portability remains limited and requires further improvement.

block size predictioncache occupancyhierarchical blocking

This work addresses the excessive overhead of conventional control mechanisms when executing multidimensional loop kernels on tightly coupled processor arrays, which severely limits performance. By leveraging the polyhedral model to represent the iteration space, the paper introduces a novel approach that expresses control conditions as unions of polyhedra, significantly simplifying control logic. Building on this formulation, the authors propose a lightweight global controller requiring hardware resources equivalent to only a single processing element, enabling zero-overhead distribution of loop control signals. Combined with bounded evaluation units and optimized control signal latency, the design reduces the number of control signals by 15–45× on the PolyBench benchmark suite, while keeping overall control-flow resource usage below 10% of the total array resources.

Control OverheadLoop ControlMultidimensional Loops

This work addresses the challenge of energy estimation for nested-loop programs on parallel processor arrays, where traditional simulation-based approaches suffer from poor scalability. To overcome this limitation, the paper proposes a symbolic polyhedral energy modeling method that, for the first time, applies symbolic polyhedral analysis to energy estimation of nested loops. By integrating loop transformation theory with array architecture modeling, the approach explicitly captures the impact of mapping and scheduling decisions on energy consumption. Experimental results demonstrate that the method achieves high-accuracy energy predictions across multiple benchmarks, with computational overhead independent of problem size, thereby significantly enhancing the scalability of design space exploration.

energy analysismapping and schedulingnested loop programs

This work addresses the challenges of latency and scalability in designing efficient concurrent primitives under high write contention in shared-memory systems. It introduces a novel approach based on a contention-resolution algorithm that transforms contention-prone hardware primitives into higher-level concurrent objects within an approximately synchronous randomized scheduling model. For the first time, the study achieves composable, low-latency concurrent primitives against an adaptive adversary, and establishes a theoretical lower bound for the space–latency tradeoff. Using only O(1) read–write registers and a single compare-and-swap (CAS) register, the construction yields—with high probability—O(log P) latency for a variety of primitives, including read–write registers, CAS, load-linked/store-conditional (LL/SC), fetch-and-increment, bounded max registers, and counters.

concurrent primitivescontention resolutionlatency

Hot Scholars

HS

Heman Shakeri

Assistant Professor, School of Data Science, University of Virginia
Learning-Based ControlData-driven Signal ProcessingComplex NetworksDynamical systems
SD

Sarah Dean

Cornell
Machine LearningOptimizationControlAlgorithmic Fairness
JF

Jianbin Fang

Associate Professor, National University of Defense Technology
High-Performance ComputingProgramming SystemsCompilersSoftware-Hardware Codesign
MF

Maryam Fazel

Moorthy Family Professor of Electrical and Computer Engineering, University of Washington
OptimizationMachine LearningControlSignal Processing
YJ

Yassir Jedra

Imperial College London
Machine LearningReinforcement LearningControl Theory