inference acceleration

Designs, builds, or analyzes systems and methods that reduce latency, increase throughput, or lower resource consumption during the inference (prediction) phase of machine learning models. Work includes model-level techniques, compiler and runtime optimizations, and hardware–software co‑design to accelerate execution while preserving acceptable accuracy and resource trade‑offs.

inferenceacceleration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

How to Keep Pushing ML Accelerator Performance? Know Your Rooflines!

May 22, 2025
MV
Marian Verhelst
🏛️ KU Leuven | imec | ETH Zurich | Universit.a di Bologna | Princeton University | EnCharge AI

To address the performance bottlenecks of machine learning (ML) accelerators under growing model sizes and stringent energy-efficiency constraints, this paper proposes the first enhanced Roofline model deeply co-designed for ML accelerator characteristics. Our method introduces *execution paradigm boundary analysis* and *energy-constrained performance upper-bound modeling*, unifying the quantification of computational intensity, memory hierarchy, and data layout effects on both performance and energy efficiency. By integrating ML workload feature extraction, architecture-level quantitative analysis, and a hardware-algorithm co-evaluation framework, we systematically identify performance bottlenecks across mainstream accelerators for diverse operators and memory layouts. Experimental results reveal synergistic optimization pathways between memory bandwidth and computational density, and clarify several open research directions. The proposed model provides both theoretical foundations and practical guidance for energy-aware ML accelerator architecture design.

Applying roofline model to optimize execution regimesEnhancing ML accelerator performance and efficiencyUnderstanding compute-memory interactions for system efficiency

Traditional AI systems rely on fixed monolithic models, which struggle to dynamically allocate resources, decompose tasks, or update knowledge in response to varying inputs, leading to degraded performance and increased costs. This work proposes the first system-level design methodology for distributed composite AI systems, formulating a design space through workflow topologies and configuration choices and identifying eight core design patterns. The framework jointly optimizes model selection and runtime parameters, enabling task decomposition, multi-model orchestration, and explicit control logic, thereby facilitating a shift from static monolithic architectures toward dynamic, composable, and adaptive ones. Evaluated across three case studies, the approach reduces latency by up to 60% and cost by up to 71%, with only a 2.5–4 percentage point drop in accuracy.

Compound AI SystemsDistributed AIModel-Centric Design

The rapid deployment of machine learning across platforms from milliwatt-class TinyML devices to large language models has made energy efficiency a primary constraint for sustainable AI. Across these scales, performance and energy are increasingly limited by data movement and memory-system behavior rather than by arithmetic throughput alone. This work reviews energy efficient software hardware codesign methods spanning edge inference and training to datacenter-scale LLM serving, covering accelerator architectures (e.g., ASIC/FPGA dataflows, processing-/compute-in-memory designs) and system-level techniques (e.g., partitioning, quantization, scheduling, and runtime adaptation). We distill common design levers and trade-offs, and highlight recurring gaps including limited cross-platform generalization, large and costly co-design search spaces, and inconsistent benchmarking across workloads and deployment settings. Finally, we outline a hierarchical decomposition perspective that maps optimization strategies to computational roles and supports incremental adaptation, offering practical guidance for building energy and carbon aware ML systems.

data movementenergy efficiencyhardware-software co-design

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

Nov 20, 2024
JF
Jared Fernandez
🏛️ Meta | Carnegie Mellon University

Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.

Assessing diminishing returns in scaling accelerators for large model trainingEvaluating parallelization strategies to minimize distributed communication overheadOptimizing hardware configuration for efficient large-scale model training

ML For Hardware Design Interpretability: Challenges and Opportunities

Apr 11, 2025
RB
Raymond Baartmans
🏛️ Oregon State University

This study addresses the poor interpretability of RTL code in hardware design by proposing the first machine learning–based automation paradigm for RTL-to-natural-language (RTL-to-NL) description generation. Methodologically, it systematically analyzes three core challenges—data scarcity, semantic gap, and domain specificity—and introduces an LLM fine-tuning framework tailored to hardware semantics, integrating RTL-structure-aware modeling with NL generation techniques. Contributions include: (1) the first systematic characterization of technical bottlenecks and evaluation dimensions for RTL-to-NL generation; (2) a hardware-aware instruction-tuning strategy coupled with verification-driven decoding; and (3) foundational theoretical and methodological support for high-fidelity RTL semantic parsing and interpretable toolchains. Experiments demonstrate significant improvements in description accuracy and functional consistency, accelerating customized AI accelerator development.

Addressing challenges in data, computation, and model developmentAutomating RTL-to-NL tasks for hardware design interpretabilityImproving efficiency of custom hardware accelerator design

Latest Papers

What's happening recently
View more

This work addresses the computational and memory bottlenecks that hinder efficient scaling in large model training. To overcome the limitations of conventional point-wise optimizations, the authors propose a throughput-centric strategy that systematically integrates multiple techniques: optimized data loading (OVERLORD), CPU memory offloading (DeepSpeed ZeRO-Offload), distributed compilation (Triton-distributed), and hardware-level dynamic voltage and frequency scaling (DVFS). This holistic approach achieves a 4.5% improvement in end-to-end training throughput, substantially reduces training costs, and enables efficient training of models significantly larger than the memory capacity of a single GPU.

computational bottlenecklarge-scale AI systemsmemory bottleneck

This work addresses the growing challenge posed by the expanding data footprint of modern applications, which renders memory systems a critical bottleneck for both performance and energy efficiency—constraints that traditional microarchitectures struggle to overcome. To this end, the paper introduces a data-driven microarchitectural design paradigm that systematically integrates lightweight machine learning with application-specific data semantic features across multiple processor components. Key contributions include a reinforcement learning–based hardware prefetcher, a perceptron-driven off-chip access predictor, a synergistic mechanism coordinating prefetching and prediction, and a predictability-aware memory access elimination technique leveraging both address and value repetition. Experimental results demonstrate that the proposed approach substantially outperforms state-of-the-art solutions, delivering significant improvements in both performance and energy efficiency.

data-aware designenergy efficiencymemory bottleneck

This work addresses the significant communication bottleneck in multi-GPU training caused by the serial execution of computation and communication. The authors propose a portable runtime mechanism that requires no modifications to vendor libraries or kernels. By dynamically controlling on-chip resource occupancy of compute kernels, elevating the scheduling priority of communication streams, and leveraging shared memory for compute footprint management and cross-GPU resource coordination, the approach effectively enables concurrent execution of computation and collective communication. Evaluated on NVIDIA A40, A100, H100, and AMD MI250X GPUs, the method reduces end-to-end training time by up to 25.5%.

communication overheadcomputation-communication overlapdistributed training

This study addresses the pressing environmental sustainability challenges posed by the high computational costs and energy consumption of large-scale AI models. It presents a systematic review of full-stack technical pathways toward greener foundation models, uniquely integrating co-optimization strategies across algorithmic and hardware layers. On the algorithmic side, it encompasses linear-complexity architectures, sparsification, and parameter-efficient fine-tuning; on the hardware side, it includes energy-efficient chips, memory-centric designs, and cross-platform deployment. The work further extends these advances to sustainability-oriented applications such as remote sensing and national infrastructure. By constructing a comprehensive roadmap spanning model design, training, and deployment, this research provides both theoretical grounding and practical guidance for developing large models that are efficient, scalable, and socially responsible.

computational costenergy consumptionenvironmental sustainability

This study systematically investigates and quantifies the critical impact of CPU bottlenecks on large language model (LLM) inference performance in multi-GPU settings, where insufficient CPU resources often lead to suboptimal GPU utilization and elevated latency—even when GPUs are not fully saturated. Through comprehensive profiling of real-world LLM serving workloads, the work identifies key CPU-induced inefficiencies, including kernel launch delays, communication stalls, and tokenization overhead. Experimental results demonstrate that moderately increasing CPU core count—without adding more GPUs—can reduce time-to-first-token latency by 1.36× to 5.40× while substantially improving system stability and responsiveness. These findings underscore the necessity of balanced CPU-GPU provisioning for efficient LLM inference deployment.

CPU bottleneckGPU underutilizationmulti-GPU LLM inference

Hot Scholars

WH

Weilin Huang

Bytedance Seed
Computer VisionDeep Learning
QG

Qiushan Guo

The University of Hong Kong; ByteDance
Deep LearningComputer Vision
WN

Weili Nie

NVIDIA Research
Machine LearningDeep LearningGenerative Models
GR

Gaël Richard

Professor, Télécom Paris, Institut polytechnique de Paris
Audio signal processingMachine listeningMusic ProcessingMusic Information Retrieval
JC

Jixiang Chen

The Hong Kong University of science and Technology
Medical reconstruction