Score
Designs and implements the runtime components and core libraries of machine learning frameworks, including automatic differentiation engines, tracing/JIT and compilation pathways, graph and optimization passes, device and memory management, and high-performance kernel dispatch. Builds and analyzes framework internals (for example JAX-like tracing, vectorized transforms and backend integration) to ensure correct, efficient, and interoperable execution across hardware backends.
The choice between PyTorch and TensorFlow remains a critical decision for AI researchers and practitioners, yet systematic, empirically grounded comparisons across usability, training/inference performance, and production deployment capabilities are lacking. Method: We conduct a comprehensive benchmarking study—including XLA, TensorRT, and other backend accelerators—analyze code complexity, evaluate cross-framework interoperability (ONNX, TorchScript, TFLite), and survey state-of-the-art literature and ecosystem tooling. Contribution/Results: Our analysis reveals fundamental paradigmatic differences: PyTorch’s dynamic computation graph excels in research agility and prototyping flexibility, whereas TensorFlow’s static graph design delivers superior end-to-end deployment maturity, multi-platform support (e.g., mobile, edge), and enterprise service integration. Computationally, both frameworks achieve comparable peak performance; however, their ecosystem roles have significantly diverged. We identify cross-framework interoperability and unified compiler-level optimization as pivotal future directions, providing evidence-based guidance for framework selection in AI development.
This work addresses the lack of a universal, flexible, and cluster-agnostic workload representation in existing distributed machine learning systems, which hinders efficient design space exploration. To overcome this limitation, the paper introduces Flint, a novel framework that leverages the intermediate representation of machine learning compilers to extract workload graphs for clusters of arbitrary scale—without requiring actual hardware execution. By decoupling workload modeling from underlying hardware specifics and validating accuracy through execution traces, Flint ensures both fidelity and portability. Experimental results demonstrate that Flint effectively enables flexible and efficient design space exploration while substantially reducing evaluation overhead.
A lack of energy-efficiency evaluation methodologies for AI inference on RISC-V architectures hinders sustainable deployment of AI workloads. Method: We conduct the first fine-grained, cross-framework energy benchmarking study—covering PyTorch, ONNX Runtime, and TensorFlow—on a production-grade 64-core SOPHON SG2042 RISC-V server, with hardware-level power monitoring and comparative analysis of XNNPACK versus OpenBLAS backends. Contribution/Results: Backend selection proves decisive for energy efficiency: enabling XNNPACK in ONNX Runtime and TensorFlow reduces average inference energy consumption by 27.3% relative to PyTorch with OpenBLAS. This work identifies critical energy bottlenecks in RISC-V AI frameworks and establishes XNNPACK as the preferred high-efficiency backend. It provides empirical evidence and actionable optimization guidance for low-carbon AI deployment across open-source ecosystems on RISC-V servers.
This study addresses the unclear usage patterns and functional demands surrounding current MLOps frameworks in open-source projects, which hinder their effective evolution. For the first time, it systematically links real-world framework adoption with user enhancement requests by analyzing GitHub dependencies, API invocations, and issue reports across eight prominent MLOps frameworks, employing qualitative coding and thematic mapping. The findings reveal that developers prefer customized integrations over out-of-the-box solutions, and that these frameworks are seldom directly embedded in GitHub Workflows, instead being primarily applied to core machine learning phases and infrastructure governance. Users most frequently request enhancements to core functionality, greater API exposure, and improved CI/CD integration, while increasingly adopting multiple frameworks in tandem.
本文提出了一种符号程序分析方法,用于机器学习内核的基本块计数,相比动态插桩技术,显著减少了运行时间。
Machine learning systems lack effective, quantifiable methods to assess how architectural design patterns impact scalability, performance, and cost—leading to subjective, evidence-deficient architecture selection. Method: This paper introduces the first quantitative architectural evaluation framework tailored for ML systems, specifically targeting CPU-based inference. It explicitly models the mappings between architectural patterns and key quality attributes: latency, throughput, and resource overhead. The framework integrates metric-driven modeling, lightweight observability analysis, and standardized benchmarking procedures. Contribution/Results: Evaluated across multiple case studies, the framework enables objective, quantitative ranking of architectural patterns—achieving up to 2.3× higher inference throughput and an average 37% improvement in CPU utilization. It further supports cost-optimized decision-making in production environments. Its core contribution is the establishment of the first rigorous, quantification-oriented paradigm for evaluating ML architectural patterns.
This work addresses the lack of efficient and reusable compilation infrastructure between high-level AI frameworks and hardware accelerators, which hinders automatic generation of high-performance code. Building upon MLIR, the authors propose a modular compiler featuring a lightweight affine analysis pipeline that integrates loop transformations, multi-level tiling, operator and attention-layer fusion, on-chip memory management, and mapping to specialized compute units. Combined with an analytical cost model and heuristic strategies, the system enables fully automated optimization from PyTorch/JAX down to hardware primitives. Evaluated on NVIDIA GPUs, the JIT-compiled code matches or exceeds the performance of Torch Inductor and XLA, with generated matrix multiplication and convolution kernels achieving parity with vendor-optimized libraries or hand-tuned kernels.
This paper addresses the challenge of automatically mapping PyTorch models to synthesizable hardware. We propose an open-source, end-to-end compilation toolchain that innovatively integrates Allo (an accelerator design language), Calyx (a hardware intermediate representation), and CIRCT (an LLVM-based hardware compilation framework). A key methodological contribution is a memory-partitioning compilation pass tailored for memory-intensive machine learning workloads, which significantly enhances data parallelism and on-chip memory efficiency. Our core contributions are threefold: (1) the first fully automated, synthesizable translation from PyTorch frontend to SystemVerilog RTL; (2) a memory optimization strategy that preserves functional correctness while substantially reducing off-chip bandwidth pressure; and (3) experimental validation showing that the generated FPGA implementations achieve throughput and resource utilization comparable to industrial-grade, closed-source tools such as Vitis HLS.
This study addresses the functional mismatches arising from the lack of hardware-level verification in compilation mappings for machine learning accelerators. To this end, it proposes BOLT, a formal verification framework that aligns loop structures with associated data layouts. Notably, BOLT introduces the first coarse-grained intrinsic verification method that operates without requiring additional auxiliary information. Furthermore, by leveraging synchronization skeleton and layout sketch template techniques, the framework achieves end-to-end correctness proofs for compiler-to-accelerator mappings. The prototype system successfully verifies the correctness of complex compilation mappings on two open-source ML accelerators, thereby providing a reliable foundation for hardware-software co-design.
本文提出SAMpLE框架,通过标准化接口将机器学习模型集成到SystemC-AMS虚拟原型中,解决了现有方法的复用性和可比性问题。
This work addresses the challenges of configuring and tuning the complex Linux kernel, where deploying machine learning (ML) models directly in kernel space is hindered by the absence of floating-point unit (FPU) support and prohibitive performance overhead. To overcome these limitations, the authors propose the first lightweight ML infrastructure tailored for the Linux kernel, enabling safe and efficient model inference without FPU usage through a cooperative kernel–user space design. The architecture comprises a kernel module, a lightweight inference agent, and a cross-space communication interface. A prototype implementation demonstrates the feasibility of this approach, and experimental results show that it introduces ML capabilities into the kernel with low overhead and high scalability, opening a new avenue for intelligent kernel optimization.