optimize data layout

Design and implement physical memory representations and transformation pipelines that map logical data structures into concrete in-memory layouts and generate alternative layouts. Build analyses and compiler or toolchain transformations that optimize those layouts to minimize copying and packing, improve alignment and efficient kernel access, and produce architecture-alapted layout variants during compilation.

optimizedatalayout

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$221K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Unified Framework for Automated Code Transformation and Pragma Insertion

May 05, 2024
SP
Stéphane Pouget
🏛️ University of California, Los Angeles | Colorado State University

In high-level synthesis (HLS), jointly optimizing code transformations, pragma insertion, and cache-blocking size selection is challenging due to tight coupling, a vast decision space, and difficulty in guaranteeing semantic correctness. Method: This paper proposes the first unified modeling framework that jointly encodes all three aspects as a single, isomorphic optimization problem—supporting “zero-transformation” decisions—and leverages HLS compiler–driven constraint derivation coupled with nonlinear programming (NLP) to automatically and correctly optimize regular loop nests. Contribution/Results: It introduces the first paradigm for co-optimizing transformations, pragmas, and blocking sizes, with built-in semantic equivalence preservation. Evaluated on multiple benchmark kernels, the approach significantly improves quality-of-results (QoR), accurately identifies cases requiring or forbidding transformations, and generates high-performance, formally verifiable optimized code.

AutomationCode ModificationSimplification

LEGO: Layout Expression for Generating One-to-one Mapping

May 12, 2025
AM
Amir Mohammad Tavakkoli
🏛️ University of Utah | University of Copenhagen

To address the optimization limitations imposed by tight coupling between data layout and computation on GPUs, this paper proposes a layout-agnostic computational abstraction paradigm: computations are first expressed in a layout-decoupled form, and hierarchical parallel index expressions are then automatically derived via layout specifications. The core contribution is the first end-to-end “layout → index expression” mapping mechanism, enabling layout-driven code generation and cross-compiler optimization exploration. We design a custom layout specification language and integrate it with MLIR, Triton, and CUDA templates to build an index derivation engine. Experimental evaluation demonstrates that the generated code achieves performance on par with hand-optimized Triton kernels. Furthermore, the approach is validated for generality and efficiency across both MLIR-based and CUDA-based compilation ecosystems.

Exploring data layouts with compiler integrationGenerating complex indexing for hierarchical GPU codeOptimizing data movement via layout-independent computation

Automatic Hardware Pragma Insertion in High-Level Synthesis: A Non-Linear Programming Approach

Apr 01, 2024
SP
Stéphane Pouget
🏛️ University of California, Los Angeles | Colorado State University

Manual pragma configuration in high-level synthesis (HLS) suffers from low efficiency and an exponentially large search space. Method: This paper proposes the first nonlinear programming (NLP)-based automated pragma insertion framework, jointly optimizing loop-level pragmas—including pipelining, function unit replication, and data caching. It innovatively models discrete pragma configurations as continuous, differentiable variables and constructs analytical performance/resource models with theoretical lower-bound guarantees, solved globally via NLP. Integrated with pragma semantic analysis and the Merlin compiler, and augmented by design-space pruning, the framework explores billion-scale configurations within seconds to minutes. Contribution/Results: Experimental evaluation shows kernel performance approaching hand-tuned implementations, resource estimation error <8%, and latency lower-bound error ≤12%.

Automatic insertion of HLS pragmasNon-linear programming for hardware synthesisOptimizing pipelining and data caching

Reimagining Disassembly Interfaces with Visualization: Combining Instruction Tracing and Control Flow with DisViz

Oct 21, 2025
SH
Shadmaan Hye
🏛️ SCI Institute | Lawrence Livermore National Laboratory

Binary disassembly analysis suffers from ambiguous source-to-instruction mapping and difficulty in jointly preserving execution order and control flow. To address this, we propose DisViz—a performance-analysis-oriented, interactive disassembly visualization tool. Its core contributions are threefold: (1) a basic-block–based instruction layout that explicitly preserves execution order while intuitively revealing control structures (e.g., loops); (2) block-level minimaps to enhance contextual awareness and navigation in large-scale disassembly; and (3) integrated instruction tracing, control-flow graph visualization, and dynamic source-code correlation, enabling bidirectional, web-based navigation between source and disassembly. An empirical evaluation with ten domain experts from diverse institutions demonstrates that DisViz significantly improves both accuracy in identifying compiler optimization behaviors and overall analysis efficiency—validating its effectiveness for understanding compilation transformations and their performance implications.

Addressing challenges in mapping binary instructions to source codeImproving developer comprehension of compiler optimizations in binariesVisualizing disassembly code with execution order and control flow

This work proposes CuTe, a novel mathematical framework for representing and manipulating hierarchical tensor layouts through a layout algebra that supports operations such as concatenation, tiling, and inversion. Modern high-performance computing and deep learning rely heavily on specialized tensor instructions whose performance and correctness are critically dependent on intricate, hardware-specific data layouts. CuTe enables unified compile-time derivation, verification, and thread mapping of these layouts, significantly simplifying GPU kernel development while supporting expressive, general-purpose tensor transformations. The framework has been integrated into production systems including the NVIDIA CUTLASS library and the CuTe DSL, effectively bridging the gap between hardware constraints and software flexibility.

data layout propagationGPU kernelshardware instructions

Latest Papers

What's happening recently
View more

This work addresses the challenge of efficiently generating vector-length-agnostic (VLA) machine learning code for scalable vector instruction sets such as Arm SVE, where unknown vector lengths at compile time hinder traditional compilers. The authors present the first end-to-end VLA support in MLIR/IREE, introducing a vector-length-aware compact data layout and unifying dynamic tiling, operator fusion, and scalable vectorization within a single compilation framework. Evaluated on Arm CPUs, the generated SVE code achieves up to 1.45× speedup over IREE’s NEON implementation, outperforms multiple frameworks in the PyTorch ecosystem, and demonstrates strong scalability with increasing vector lengths in simulation, effectively balancing performance and hardware portability.

compiler code generationdata layoutML compilation

This work addresses the performance limitations of traditional pointer analysis by proposing a decoupled acceleration paradigm. Instead of tightly coupling simplification rules with the analysis itself—a common drawback of existing offline approaches—the method applies general-purpose, semantics-preserving compiler optimizations to the intermediate representation (IR) prior to analysis. This modular, analysis-agnostic strategy enhances efficiency without requiring modifications to the pointer analysis algorithm, thereby supporting seamless integration with diverse analyses. Empirical evaluation across multiple benchmark programs and three mainstream pointer analyses demonstrates speedups of up to 3.14× and memory reductions of up to 1.94×, all while largely preserving precision.

compiler optimizationintermediate representationoffline simplification

This work addresses the challenge of energy estimation for nested-loop programs on parallel processor arrays, where traditional simulation-based approaches suffer from poor scalability. To overcome this limitation, the paper proposes a symbolic polyhedral energy modeling method that, for the first time, applies symbolic polyhedral analysis to energy estimation of nested loops. By integrating loop transformation theory with array architecture modeling, the approach explicitly captures the impact of mapping and scheduling decisions on energy consumption. Experimental results demonstrate that the method achieves high-accuracy energy predictions across multiple benchmarks, with computational overhead independent of problem size, thereby significantly enhancing the scalability of design space exploration.

energy analysismapping and schedulingnested loop programs

Hot Scholars

MS

Mohammad Sadrosadati

Senior Researcher and Lecturer, ETH Zürich
Heterogeneous ComputingProcessing-In-MemoryMemory SystemsInterconnection Networks
CF

Chao Fang

Shanghai Qi Zhi Institute
efficient MLAI acceleratorhardware-software co-designprecision-scalable computing
CG

Christina Giannoula

Postdoctoral Researcher, University of Toronto
Computer ArchitectureComputer SystemsProcessing-In-MemoryMachine Learning
FL

Fangxin Liu

Shanghai Jiao Tong University
In-memory Computing、Brian-inspired Neuromorphic Computing