implement operator fusion

Designs, implements, or analyzes methods that consolidate multiple computation or data-processing steps into a single fused operator or representation to reduce dispatch overhead, memory traffic, duplicated hardware resources, and improve throughput. This includes kernel-level operator fusion (for example fused matmul+bias+activation), hardware-aware and operator-consolidation strategies that map several functions onto shared hardware with minimal area, and techniques for combining multiple data, signal, or feature sources into a single fused representation.

implementoperatorfusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.39
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$219K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators

Jun 27, 2025
ZZ
Zheng Zhang
🏛️ Wuhan University | NVIDIA | University of Macau

Traditional fusion approaches for memory-constrained, compute-intensive dynamic tensor operator chains suffer from narrow search spaces, redundant memory accesses, and high tuning overhead. Method: This paper proposes MCFuser—a framework that (i) formally defines such operator chains; (ii) constructs a complete fusion strategy search space using high-dimensional tiling expressions; and (iii) integrates DAG-driven memory access optimization with an analytical performance model-guided heuristic search to enable efficient pruning and automatic kernel generation. Contribution/Results: Evaluated on NVIDIA A100 and RTX 3080 GPUs, MCFuser achieves up to 5.9× higher kernel performance and reduces tuning time by 70× compared to state-of-the-art compilers (e.g., Ansor). It significantly improves GPU data locality and alleviates memory bandwidth bottlenecks.

Dynamic tensor dimensions causing memory-bound operators requiring fused kernelsFusion of multiple compute-intensive operators hindered by computation saturationLimited fusion strategy search spaces and redundant memory access degrading performance

The Fused Kernel Library: A C++ API to Develop Highly-Efficient GPU Libraries

Aug 09, 2025
OA
Oscar Amoros
🏛️ Universitat Politecnica de Catalunya | NVIDIA | Barcelona Supercomputing Center (BSC)

Existing GPU libraries rely on manually compiled fused kernels, hindering flexible support for horizontal and vertical fusion (HF/VF), resulting in suboptimal on-chip SRAM utilization and frequent spilling of intermediate data to global memory. This work proposes an automatic kernel fusion framework based on C++17 template metaprogramming: users declaratively define composable GPU function components, and the system statically infers and generates optimal fused kernels at compile time—natively supporting both HF and VF while eliminating unnecessary global memory transfers. The approach requires no precompiled templates or hand-written fusion code, drastically reducing development overhead and enhancing programming flexibility and hardware resource efficiency. An open-source implementation demonstrates 2×–1000× speedups over state-of-the-art GPU libraries across diverse benchmarks, achieving, for the first time, a principled unification of high performance and high programmability.

Enables automatic kernel fusion for arbitrary GPU function combinationsMaximizes GPU resource utilization and on-chip memory efficiencyReduces development costs by eliminating manual fused kernel creation

This work addresses the inefficiency and suboptimal energy consumption of data aggregation operations on heterogeneous hardware platforms. To bridge this gap, the authors propose a hybrid hardware acceleration framework that synergistically combines unified abstractions with platform-specific optimizations, effectively balancing programmability, portability, and architectural specialization across CPUs, GPUs, and FPGAs. By introducing a common abstraction layer while incorporating tailored optimization strategies for each hardware target, the approach achieves significant improvements in both performance and energy efficiency across all three mainstream architectures. The evaluation demonstrates consistent gains not only in device-level computation but also in end-to-end processing metrics, thereby establishing an effective trade-off between generality and high performance for data aggregation workloads.

aggregationdata analyticsFPGA

This work addresses the limitations of existing deep learning compilers, which optimize compute-intensive and memory-intensive operators in isolation, leading to rigid fusion boundaries that hinder cross-operator optimization and on-chip data reuse. To overcome this, the paper proposes a novel operator fusion strategy that, for the first time, enables synergistic fusion of both types of subgraphs. The approach automatically integrates complex, dependency-rich memory-intensive subgraphs with compute-intensive operators—such as back-to-back GEMMs—into high-performance GPU kernels. Built upon automated GPU compilation techniques, the method lowers high-level tensor programs into optimized fused kernels and employs concurrent scheduling to hide memory latency. Experimental results demonstrate that the generated kernels outperform those produced by TorchInductor on complex workloads like post-norm transformers and support a broader range of fusion patterns.

compute-intensive operatorsGPU compilationheterogeneous computation graphs

Approximate Computing Survey, Part I: Terminology and Software & Hardware Approximation Techniques

Jul 20, 2023
VL
Vasileios Leon
🏛️ National Technical University of Athens | New York University | Villanova University

Emerging edge and cloud AI applications demand high-energy-efficiency computing, yet conventional embedded and datacenter architectures struggle to simultaneously achieve high performance and energy efficiency. Method: This work systematically surveys 15 years of approximate computing research, introducing the first full-stack taxonomy—spanning programs, compilers, circuits, accelerators, and memory—along with rigorously defined core terminology and design principles; it further proposes a unified evaluation framework for quantitative, cross-layer trade-off analysis between performance and power consumption. Contribution/Results: The study delivers the first authoritative survey on approximate computing (Part I), addressing a critical gap in systematic, domain-wide reviews. By establishing foundational taxonomies and evaluation methodologies, it provides both theoretical grounding and practical guidance for algorithm–architecture co-optimization, thereby advancing energy-efficient computing for AI workloads.

Address high-performance demands in multimedia and machine learning applications.Explore and classify software and hardware approximation techniques.Improve power efficiency in computing systems for embedded and data centers.

Latest Papers

What's happening recently
View more

This work addresses the challenge of efficiently deploying Mamba models on modern hardware, which is hindered by complex data dependencies leading to excessive off-chip memory accesses. To tackle this, the authors introduce a cascade-of-Einsums abstraction to systematically characterize Mamba’s computational structure and devise cross-Einsum fusion strategies that minimize data movement. Building upon this foundation, they design Mambalaya, a reconfigurable accelerator that pioneers the use of an extended Einsum framework to uniformly fuse multiple operators within Mamba. Experimental results demonstrate that Mambalaya achieves a 4.9× speedup over MARCA during the prefill phase and a 1.9× speedup in the generation phase; moreover, in prefill-dominated scenarios, it delivers 1.5× higher performance than existing fused accelerators.

EinsumfusionMamba

This work addresses the computational intensity, low energy efficiency, and insufficient on-chip data reuse in large language model (LLM) inference by proposing a fusion-driven compute-in-memory (CIM) architecture. It co-designs the attention mechanism by synergistically fusing QKᵀ and PV computations, innovatively integrating input-side (IP-CIM) and output-side (OP-CIM) compute-in-memory paradigms, and introducing a QO-stationary dataflow to maximize on-chip data reuse. Additionally, a pattern-aware online Softmax mechanism is incorporated to substantially reduce the overhead of nonlinear operations. Experimental evaluation on the LLaMA-3 model demonstrates that the proposed architecture achieves up to 1.98× speedup and 3.86× energy savings, attaining a system energy efficiency of 29.4 TOPS/W.

attention mechanismcompute-in-memoryenergy efficiency

Existing edge AI inference systems are constrained by model-level mapping strategies, which hinder efficient utilization of heterogeneous computing resources to accommodate diverse operator characteristics. This work proposes the first unified operator-level scheduling framework that dynamically assigns each operator to the optimal processing unit (CPU/GPU/NPU) based on empirical performance profiling. By constructing a weighted execution graph and solving a shortest-path problem, the framework enables latency- or energy-efficiency-oriented scheduling. It transcends conventional limitations by uniformly supporting sequential execution, intra-model parallelism, and multi-model concurrency, all without relying on model-specific heuristics, thus achieving model-agnostic applicability. Experiments on an Intel Core Ultra SoC demonstrate up to 1.60× speedup with intra-model parallelism, a geometric mean acceleration of 3.42× for concurrent multi-model execution, and an average energy saving of 48.2% under energy-efficient scheduling.

edge AIheterogeneous edge inferencemodel heterogeneity

Existing mappers struggle to find optimal dataflow mappings with operator fusion for tensor algebra accelerators within a reasonable time, as their search space grows exponentially with the number of computation steps. This work proposes the Fast and Fusiest Mapper (FFM), the first approach capable of efficiently searching the complete fused mapping space for optimality. FFM introduces a fusion-aware pruning strategy that eliminates suboptimal partial mappings early, combined with accurate performance modeling and partial-mapping stitching techniques to drastically reduce the search space. Evaluated on Transformer workloads, FFM achieves over 1,000× speedup compared to the state-of-the-art while exhibiting near-linear runtime scaling, effectively overcoming the exponential complexity barrier inherent in fused mapping exploration.

accelerator modelingfusionmapper

This work systematically investigates the trade-offs between performance and programmability in CPU-GPU cooperative scheduling across discrete and unified memory architectures, with a focus on sparse conjugate gradient computations. Evaluations are conducted on both the NVIDIA GH200 Superchip—a platform featuring a unified memory architecture—and the discrete H100 PCIe system, comparing three memory management paradigms: explicit data copies, managed memory, and mapped memory. The study reveals that the GH200’s fused architecture substantially enhances the practicality of managed memory, enabling diverse hybrid task-partitioning strategies to achieve both high performance and programming simplicity. These findings underscore the significant impact of underlying memory architecture on the efficacy of cooperative scheduling approaches.

coschedulingCPU-GPUintegrated GPU

Hot Scholars

MG

Minyi Guo

IEEE Fellow, Chair Professor, Shanghai Jiao Tong University
Parallel ComputingCompiler OptimizationCloud ComputingNetworking
MV

Marian Verhelst

Micas - ESAT - KU Leuven, Belgium
Low-energy chip designsensor fusionmachine learningcross-layer optimization
JL

Jingwen Leng

Professor, Shanghai Jiao Tong University
Computer Architecture