deep learning framework engineering

Designs, implements, and analyzes the internals, runtime systems, compiler passes, and integration layers of deep learning frameworks, with a focus on components such as execution graph compilation, backend adapters, and memory/kernel scheduling. Builds and evaluates optimizations for framework execution (e.g., torch.compile and other PyTorch compiler features), profiles and tunes runtime behavior, and integrates frameworks with hardware and external toolchains to improve throughput, latency, and resource usage.

deeplearningframeworkengineering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.49
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$215K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

A High-Level Compiler Integration Approach for Deep Learning Accelerators Supporting Abstraction and Optimization

Jul 07, 2025
SA
Samira Ahmadifarsani
🏛️ Technical University of Munich | TU Wien

Integrating custom hardware accelerators—particularly GEMM-based ones—into mainstream ML compilers remains challenging due to tight coupling between accelerator-specific optimizations and compiler internals. Method: This paper proposes a high-level, low-intrusion integration methodology for TVM that abstracts hardware scheduling interfaces to decouple accelerator characteristics from compiler implementation. Leveraging the CoSA design-space exploration framework, it automates hardware-aware scheduling optimizations—including tensor tiling, non-uniform mapping, and double buffering—without modifying TVM’s core infrastructure. Contribution/Results: Evaluated on the Gemmini accelerator, the approach achieves performance on par with hand-optimized toolchains while significantly improving developer productivity and cross-model/cross-architecture portability. The abstraction enables seamless reuse of scheduling policies across diverse accelerator microarchitectures and neural network workloads, reducing integration effort from weeks to hours.

Automating efficient tensor scheduling for deep learning accelerators is challengingExisting frameworks require deep compiler knowledge and offer partial solutionsIntegrating custom accelerators into ML compilers is complex

This study addresses the lack of automated Processing-in-Memory (PIM) offloading support in deep learning frameworks and the associated data movement bottlenecks by proposing a compiler-based automated host-PIM placement strategy. Leveraging MLIR and PyTorch compiler technologies, this method overcomes the limitations of fixed operator lists by dynamically profiling the computational intensity and memory access characteristics of loop nests generated through progressive loop order lowering. This enables profile-guided optimization (PGO)-driven automatic offloading decisions. Evaluated across diverse PIM configurations, the proposed approach achieves speedups of 5.1× and 3.6× over CPU-only execution for GPT-J-6B and LLaMA-7B, respectively.

Compiler OffloadingData Movement BottleneckDeep Learning

Exploring energy consumption of AI frameworks on a 64-core RV64 Server CPU

Apr 03, 2025
GM
Giulio Malenza
🏛️ University of Torino

A lack of energy-efficiency evaluation methodologies for AI inference on RISC-V architectures hinders sustainable deployment of AI workloads. Method: We conduct the first fine-grained, cross-framework energy benchmarking study—covering PyTorch, ONNX Runtime, and TensorFlow—on a production-grade 64-core SOPHON SG2042 RISC-V server, with hardware-level power monitoring and comparative analysis of XNNPACK versus OpenBLAS backends. Contribution/Results: Backend selection proves decisive for energy efficiency: enabling XNNPACK in ONNX Runtime and TensorFlow reduces average inference energy consumption by 27.3% relative to PyTorch with OpenBLAS. This work identifies critical energy bottlenecks in RISC-V AI frameworks and establishes XNNPACK as the preferred high-efficiency backend. It provides empirical evidence and actionable optimization guidance for low-carbon AI deployment across open-source ecosystems on RISC-V servers.

Analyzing energy consumption of AI frameworks on RISC-V CPUComparing energy efficiency of PyTorch, ONNX Runtime, TensorFlowEvaluating impact of back-end choices on AI framework energy use

From PyTorch to Calyx: An Open-Source Compiler Toolchain for ML Accelerators

Dec 05, 2025
JX
Jiahan Xie
🏛️ University of California, Santa Cruz | Cornell University

This paper addresses the challenge of automatically mapping PyTorch models to synthesizable hardware. We propose an open-source, end-to-end compilation toolchain that innovatively integrates Allo (an accelerator design language), Calyx (a hardware intermediate representation), and CIRCT (an LLVM-based hardware compilation framework). A key methodological contribution is a memory-partitioning compilation pass tailored for memory-intensive machine learning workloads, which significantly enhances data parallelism and on-chip memory efficiency. Our core contributions are threefold: (1) the first fully automated, synthesizable translation from PyTorch frontend to SystemVerilog RTL; (2) a memory optimization strategy that preserves functional correctness while substantially reducing off-chip bandwidth pressure; and (3) experimental validation showing that the generated FPGA implementations achieve throughput and resource utilization comparable to industrial-grade, closed-source tools such as Vitis HLS.

Compiles PyTorch models to SystemVerilog for acceleratorsEnables memory partitioning for parallelism in ML workloadsGenerates optimized FPGA designs comparable to closed-source tools

This study addresses a critical yet underexplored class of defects—referred to as fBugs—in the frontends of deep learning compilers during the conversion of programs into graph-based intermediate representations. Focusing on TorchDynamo, the default frontend of PyTorch 2, this work presents the first systematic empirical investigation of fBugs by leveraging a domain-knowledge-enhanced large language model to analyze 123 real-world bugs. The authors establish a comprehensive taxonomy encompassing seven root-cause categories and fifteen subcategories, and develop root-cause-aware test cases. Moving beyond conventional black-box or low-level API–centric approaches, their methodology successfully uncovers 23 previously unknown fBugs—15 of which have been confirmed—spanning eight subcategories, thereby substantially improving the robustness of compiler frontends.

deep learning compilerempirical studyfrontend bugs

Latest Papers

What's happening recently
View more

This work addresses the lack of a universal, flexible, and cluster-agnostic workload representation in existing distributed machine learning systems, which hinders efficient design space exploration. To overcome this limitation, the paper introduces Flint, a novel framework that leverages the intermediate representation of machine learning compilers to extract workload graphs for clusters of arbitrary scale—without requiring actual hardware execution. By decoupling workload modeling from underlying hardware specifics and validating accuracy through execution traces, Flint ensures both fidelity and portability. Experimental results demonstrate that Flint effectively enables flexible and efficient design space exploration while substantially reducing evaluation overhead.

compiler intermediate representationdesign space explorationdistributed machine learning

This work proposes ExecuTorch, the first end-to-end deployment framework natively integrated with the PyTorch ecosystem, addressing the fragmentation commonly encountered in edge AI deployment. By introducing an extensible backend abstraction, quantization-aware optimizations, and a unified model serialization format, ExecuTorch preserves the original model semantics while seamlessly targeting heterogeneous hardware—from microcontrollers to specialized accelerators—without sacrificing low latency or offline execution capabilities. The framework bridges the gap between research and production workflows, enabling consistent development and efficient deployment across a broad spectrum of devices, ranging from wearables to compute clusters, thereby significantly enhancing both deployment efficiency and cross-platform consistency.

edge AIhardware heterogeneitymodel deployment

This work addresses silent correctness bugs in the PyTorch compiler (torch.compile), which produce erroneous model outputs without raising errors, thereby posing a critical threat to the reliability of large language models. The study presents the first systematic empirical analysis of the characteristics of such defects and introduces AlignGuard, the first targeted detection method designed specifically for this problem. AlignGuard integrates fuzz testing with large language model–guided test case mutation, augmented by a rigorous correctness validation mechanism. The approach successfully uncovered 23 previously unknown bugs, all of which have been acknowledged or fixed by the PyTorch team, including 14 classified as high-priority, substantially enhancing the compiler’s trustworthiness.

correctness bugsLLM reliabilityPyTorch compiler

This work addresses the challenge of efficiently detecting security vulnerabilities in deep learning frameworks, which stems from their multilingual architecture and complex tensor states that hinder traditional static analysis. The authors propose Phoenix, the first large language model (LLM)-based static analysis approach, which introduces a novel Semantic Bridge Intermediate Representation (SBIR) to model cross-language tensor flows. Phoenix employs a multi-agent collaborative workflow that integrates historical patches, CWE rules, code symbol retrieval, and SBIR generation to enable precise vulnerability detection without runtime execution, facilitating accurate tensor semantic propagation analysis. Evaluated on PyTorch, Phoenix successfully identified 31 previously unknown real-world vulnerabilities across Intel CPU, NVIDIA CUDA, and Apple MPS backends, with 20 already patched and merged upstream.

deep learning frameworksmultilingual architecturessecurity bugs

Industrial-scale recommendation and ranking models feature highly complex and continuously evolving architectures, rendering traditional optimization approaches—based on manual intervention or module-level rules—difficult to scale. This work proposes the first extensible and customizable operator-level automatic transformation framework integrated into PyTorch 2.x. By leveraging FX intermediate representation, the PT2 compiler, predefined pattern matching, and a greedy search algorithm, the framework achieves general-purpose model optimizations while strictly preserving computational semantics. Evaluated on real-world industrial recommendation models, the approach delivers up to 63% inference speedup, a 6% reduction in peak memory usage, and over 400 seconds of compilation time savings. The implementation has been open-sourced as part of PyTorch 2.x.

deep learninggraph optimizationmodel transformation

Hot Scholars

RB

Rebekka Burkholz

CISPA Helmholtz Center for Information Security
Machine LearningDeep Learning EfficiencyComplex NetworksCascades
YZ

Yunfei Zhao

Peking University
intelligent programcode generationcode representation
DJ

Dag Johansen

Professor Informatikk UiT - The Arctic University of Norway
large-scale systems