ml framework runtime and library development

Designs and implements the runtime components and core libraries of machine learning frameworks, including automatic differentiation engines, tracing/JIT and compilation pathways, graph and optimization passes, device and memory management, and high-performance kernel dispatch. Builds and analyzes framework internals (for example JAX-like tracing, vectorized transforms and backend integration) to ensure correct, efficient, and interoperable execution across hardware backends.

mlframeworkruntimeand

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$228K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of a universal, flexible, and cluster-agnostic workload representation in existing distributed machine learning systems, which hinders efficient design space exploration. To overcome this limitation, the paper introduces Flint, a novel framework that leverages the intermediate representation of machine learning compilers to extract workload graphs for clusters of arbitrary scale—without requiring actual hardware execution. By decoupling workload modeling from underlying hardware specifics and validating accuracy through execution traces, Flint ensures both fidelity and portability. Experimental results demonstrate that Flint effectively enables flexible and efficient design space exploration while substantially reducing evaluation overhead.

compiler intermediate representationdesign space explorationdistributed machine learning

Exploring energy consumption of AI frameworks on a 64-core RV64 Server CPU

Apr 03, 2025
GM
Giulio Malenza
🏛️ University of Torino

A lack of energy-efficiency evaluation methodologies for AI inference on RISC-V architectures hinders sustainable deployment of AI workloads. Method: We conduct the first fine-grained, cross-framework energy benchmarking study—covering PyTorch, ONNX Runtime, and TensorFlow—on a production-grade 64-core SOPHON SG2042 RISC-V server, with hardware-level power monitoring and comparative analysis of XNNPACK versus OpenBLAS backends. Contribution/Results: Backend selection proves decisive for energy efficiency: enabling XNNPACK in ONNX Runtime and TensorFlow reduces average inference energy consumption by 27.3% relative to PyTorch with OpenBLAS. This work identifies critical energy bottlenecks in RISC-V AI frameworks and establishes XNNPACK as the preferred high-efficiency backend. It provides empirical evidence and actionable optimization guidance for low-carbon AI deployment across open-source ecosystems on RISC-V servers.

Analyzing energy consumption of AI frameworks on RISC-V CPUComparing energy efficiency of PyTorch, ONNX Runtime, TensorFlowEvaluating impact of back-end choices on AI framework energy use

This study addresses the unclear usage patterns and functional demands surrounding current MLOps frameworks in open-source projects, which hinder their effective evolution. For the first time, it systematically links real-world framework adoption with user enhancement requests by analyzing GitHub dependencies, API invocations, and issue reports across eight prominent MLOps frameworks, employing qualitative coding and thematic mapping. The findings reveal that developers prefer customized integrations over out-of-the-box solutions, and that these frameworks are seldom directly embedded in GitHub Workflows, instead being primarily applied to core machine learning phases and infrastructure governance. Users most frequently request enhancements to core functionality, greater API exposure, and improved CI/CD integration, while increasingly adopting multiple frameworks in tandem.

empirical studyfeature requestsframework usage

A quantitative framework for evaluating architectural patterns in ML systems

Jan 20, 2025
SE
Simeon Emanuilov
🏛️ Sofia University 'St. Kliment Ohridski'

Machine learning systems lack effective, quantifiable methods to assess how architectural design patterns impact scalability, performance, and cost—leading to subjective, evidence-deficient architecture selection. Method: This paper introduces the first quantitative architectural evaluation framework tailored for ML systems, specifically targeting CPU-based inference. It explicitly models the mappings between architectural patterns and key quality attributes: latency, throughput, and resource overhead. The framework integrates metric-driven modeling, lightweight observability analysis, and standardized benchmarking procedures. Contribution/Results: Evaluated across multiple case studies, the framework enables objective, quantitative ranking of architectural patterns—achieving up to 2.3× higher inference throughput and an average 37% improvement in CPU utilization. It further supports cost-optimized decision-making in production environments. Its core contribution is the establishment of the first rigorous, quantification-oriented paradigm for evaluating ML architectural patterns.

Design Pattern EvaluationMachine Learning SystemsScalability and Efficiency

Latest Papers

What's happening recently
View more

This work addresses the lack of efficient and reusable compilation infrastructure between high-level AI frameworks and hardware accelerators, which hinders automatic generation of high-performance code. Building upon MLIR, the authors propose a modular compiler featuring a lightweight affine analysis pipeline that integrates loop transformations, multi-level tiling, operator and attention-layer fusion, on-chip memory management, and mapping to specialized compute units. Combined with an analytical cost model and heuristic strategies, the system enables fully automated optimization from PyTorch/JAX down to hardware primitives. Evaluated on NVIDIA GPUs, the JIT-compiled code matches or exceeds the performance of Torch Inductor and XLA, with generated matrix multiplication and convolution kernels achieving parity with vendor-optimized libraries or hand-tuned kernels.

AI chipsAI programming frameworksautomatic code generation

From PyTorch to Calyx: An Open-Source Compiler Toolchain for ML Accelerators

Dec 05, 2025
JX
Jiahan Xie
🏛️ University of California, Santa Cruz | Cornell University

This paper addresses the challenge of automatically mapping PyTorch models to synthesizable hardware. We propose an open-source, end-to-end compilation toolchain that innovatively integrates Allo (an accelerator design language), Calyx (a hardware intermediate representation), and CIRCT (an LLVM-based hardware compilation framework). A key methodological contribution is a memory-partitioning compilation pass tailored for memory-intensive machine learning workloads, which significantly enhances data parallelism and on-chip memory efficiency. Our core contributions are threefold: (1) the first fully automated, synthesizable translation from PyTorch frontend to SystemVerilog RTL; (2) a memory optimization strategy that preserves functional correctness while substantially reducing off-chip bandwidth pressure; and (3) experimental validation showing that the generated FPGA implementations achieve throughput and resource utilization comparable to industrial-grade, closed-source tools such as Vitis HLS.

Compiles PyTorch models to SystemVerilog for acceleratorsEnables memory partitioning for parallelism in ML workloadsGenerates optimized FPGA designs comparable to closed-source tools

This study addresses the functional mismatches arising from the lack of hardware-level verification in compilation mappings for machine learning accelerators. To this end, it proposes BOLT, a formal verification framework that aligns loop structures with associated data layouts. Notably, BOLT introduces the first coarse-grained intrinsic verification method that operates without requiring additional auxiliary information. Furthermore, by leveraging synchronization skeleton and layout sketch template techniques, the framework achieves end-to-end correctness proofs for compiler-to-accelerator mappings. The prototype system successfully verifies the correctness of complex compilation mappings on two open-source ML accelerators, thereby providing a reliable foundation for hardware-software co-design.

Compiler-to-Accelerator MappingsFormal VerificationFunctional Equivalence

This work addresses the challenges of configuring and tuning the complex Linux kernel, where deploying machine learning (ML) models directly in kernel space is hindered by the absence of floating-point unit (FPU) support and prohibitive performance overhead. To overcome these limitations, the authors propose the first lightweight ML infrastructure tailored for the Linux kernel, enabling safe and efficient model inference without FPU usage through a cooperative kernel–user space design. The architecture comprises a kernel module, a lightweight inference agent, and a cross-space communication interface. A prototype implementation demonstrates the feasibility of this approach, and experimental results show that it introduces ML capabilities into the kernel with low overhead and high scalability, opening a new avenue for intelligent kernel optimization.

floating-point operationskernel spaceLinux kernel

Hot Scholars

SD

Srinivas Devadas

Edwin Sibley Webster Professor of Electrical Engineering and Computer Science, MIT
Applied cryptographyComputer securityComputer architectureComputer-Aided Design
WC

Wilka Carvalho

Harvard University
cognitive sciencereinforcement learningdeep learning
OP

Olivier Peltre

InstaDeep
mathematical physicstopology and geometryhigh dimensional statisticscomputer science