Score
Designs, builds, or analyzes systems and methods that reduce latency, increase throughput, or lower resource consumption during the inference (prediction) phase of machine learning models. Work includes model-level techniques, compiler and runtime optimizations, and hardware–software co‑design to accelerate execution while preserving acceptable accuracy and resource trade‑offs.
Designing deep learning accelerators for heterogeneous HPC and edge platforms faces key challenges including insufficient parallelism exploitation and excessive data movement overhead. This paper systematically surveys accelerator design methodologies, covering hardware-software co-design, high-level synthesis, domain-specific compilers (e.g., TVM, Halide), design space exploration, and cycle-accurate modeling and simulation. We propose, for the first time, a unified multi-dimensional classification framework that distills two fundamental principles: “minimizing data movement” and “maximizing parallelism.” The survey bridges the gap between architectural overviews and implementation-oriented methodologies, explicitly identifying emerging directions such as approximate computing integrated with reconfigurability. Our work provides both a methodological foundation and practical guidance for developing efficient, scalable AI accelerators—enabling principled design decisions across diverse heterogeneous computing ecosystems.
To address the performance bottlenecks of machine learning (ML) accelerators under growing model sizes and stringent energy-efficiency constraints, this paper proposes the first enhanced Roofline model deeply co-designed for ML accelerator characteristics. Our method introduces *execution paradigm boundary analysis* and *energy-constrained performance upper-bound modeling*, unifying the quantification of computational intensity, memory hierarchy, and data layout effects on both performance and energy efficiency. By integrating ML workload feature extraction, architecture-level quantitative analysis, and a hardware-algorithm co-evaluation framework, we systematically identify performance bottlenecks across mainstream accelerators for diverse operators and memory layouts. Experimental results reveal synergistic optimization pathways between memory bandwidth and computational density, and clarify several open research directions. The proposed model provides both theoretical foundations and practical guidance for energy-aware ML accelerator architecture design.
Traditional AI systems rely on fixed monolithic models, which struggle to dynamically allocate resources, decompose tasks, or update knowledge in response to varying inputs, leading to degraded performance and increased costs. This work proposes the first system-level design methodology for distributed composite AI systems, formulating a design space through workflow topologies and configuration choices and identifying eight core design patterns. The framework jointly optimizes model selection and runtime parameters, enabling task decomposition, multi-model orchestration, and explicit control logic, thereby facilitating a shift from static monolithic architectures toward dynamic, composable, and adaptive ones. Evaluated across three case studies, the approach reduces latency by up to 60% and cost by up to 71%, with only a 2.5–4 percentage point drop in accuracy.
The rapid deployment of machine learning across platforms from milliwatt-class TinyML devices to large language models has made energy efficiency a primary constraint for sustainable AI. Across these scales, performance and energy are increasingly limited by data movement and memory-system behavior rather than by arithmetic throughput alone. This work reviews energy efficient software hardware codesign methods spanning edge inference and training to datacenter-scale LLM serving, covering accelerator architectures (e.g., ASIC/FPGA dataflows, processing-/compute-in-memory designs) and system-level techniques (e.g., partitioning, quantization, scheduling, and runtime adaptation). We distill common design levers and trade-offs, and highlight recurring gaps including limited cross-platform generalization, large and costly co-design search spaces, and inconsistent benchmarking across workloads and deployment settings. Finally, we outline a hierarchical decomposition perspective that maps optimization strategies to computational roles and supports incremental adaptation, offering practical guidance for building energy and carbon aware ML systems.
Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.
This study addresses the poor interpretability of RTL code in hardware design by proposing the first machine learning–based automation paradigm for RTL-to-natural-language (RTL-to-NL) description generation. Methodologically, it systematically analyzes three core challenges—data scarcity, semantic gap, and domain specificity—and introduces an LLM fine-tuning framework tailored to hardware semantics, integrating RTL-structure-aware modeling with NL generation techniques. Contributions include: (1) the first systematic characterization of technical bottlenecks and evaluation dimensions for RTL-to-NL generation; (2) a hardware-aware instruction-tuning strategy coupled with verification-driven decoding; and (3) foundational theoretical and methodological support for high-fidelity RTL semantic parsing and interpretable toolchains. Experiments demonstrate significant improvements in description accuracy and functional consistency, accelerating customized AI accelerator development.
This work addresses the computational and memory bottlenecks that hinder efficient scaling in large model training. To overcome the limitations of conventional point-wise optimizations, the authors propose a throughput-centric strategy that systematically integrates multiple techniques: optimized data loading (OVERLORD), CPU memory offloading (DeepSpeed ZeRO-Offload), distributed compilation (Triton-distributed), and hardware-level dynamic voltage and frequency scaling (DVFS). This holistic approach achieves a 4.5% improvement in end-to-end training throughput, substantially reduces training costs, and enables efficient training of models significantly larger than the memory capacity of a single GPU.
This work addresses the growing challenge posed by the expanding data footprint of modern applications, which renders memory systems a critical bottleneck for both performance and energy efficiency—constraints that traditional microarchitectures struggle to overcome. To this end, the paper introduces a data-driven microarchitectural design paradigm that systematically integrates lightweight machine learning with application-specific data semantic features across multiple processor components. Key contributions include a reinforcement learning–based hardware prefetcher, a perceptron-driven off-chip access predictor, a synergistic mechanism coordinating prefetching and prediction, and a predictability-aware memory access elimination technique leveraging both address and value repetition. Experimental results demonstrate that the proposed approach substantially outperforms state-of-the-art solutions, delivering significant improvements in both performance and energy efficiency.
This work addresses the significant communication bottleneck in multi-GPU training caused by the serial execution of computation and communication. The authors propose a portable runtime mechanism that requires no modifications to vendor libraries or kernels. By dynamically controlling on-chip resource occupancy of compute kernels, elevating the scheduling priority of communication streams, and leveraging shared memory for compute footprint management and cross-GPU resource coordination, the approach effectively enables concurrent execution of computation and collective communication. Evaluated on NVIDIA A40, A100, H100, and AMD MI250X GPUs, the method reduces end-to-end training time by up to 25.5%.
This study addresses the pressing environmental sustainability challenges posed by the high computational costs and energy consumption of large-scale AI models. It presents a systematic review of full-stack technical pathways toward greener foundation models, uniquely integrating co-optimization strategies across algorithmic and hardware layers. On the algorithmic side, it encompasses linear-complexity architectures, sparsification, and parameter-efficient fine-tuning; on the hardware side, it includes energy-efficient chips, memory-centric designs, and cross-platform deployment. The work further extends these advances to sustainability-oriented applications such as remote sensing and national infrastructure. By constructing a comprehensive roadmap spanning model design, training, and deployment, this research provides both theoretical grounding and practical guidance for developing large models that are efficient, scalable, and socially responsible.
This study systematically investigates and quantifies the critical impact of CPU bottlenecks on large language model (LLM) inference performance in multi-GPU settings, where insufficient CPU resources often lead to suboptimal GPU utilization and elevated latency—even when GPUs are not fully saturated. Through comprehensive profiling of real-world LLM serving workloads, the work identifies key CPU-induced inefficiencies, including kernel launch delays, communication stalls, and tokenization overhead. Experimental results demonstrate that moderately increasing CPU core count—without adding more GPUs—can reduce time-to-first-token latency by 1.36× to 5.40× while substantially improving system stability and responsiveness. These findings underscore the necessity of balanced CPU-GPU provisioning for efficient LLM inference deployment.