ai accelerators

Designs, builds, and evaluates hardware and hardware–software systems that accelerate AI/ML computation, including custom accelerators, FPGA/GPU-based designs, memory and interconnect architectures, compilers and runtimes, and support for numerical formats and quantization. Analyzes performance, latency, throughput, area and energy efficiency, and the trade-offs between hardware, software, and model changes to optimize model execution.

aiaccelerators

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.54
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$189K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

The computational demands of AI models—particularly deep neural networks (DNNs)—continue to outpace the capabilities of conventional architectures. Method: This paper systematically analyzes design principles and performance trade-offs across mainstream AI accelerators—including GPUs, ASICs, and FPGAs—employing architectural analysis, fine-grained performance modeling, and multi-dimensional benchmarking to investigate key techniques: dataflow optimization, memory hierarchy restructuring, sparsity exploitation, and low-precision quantization. Contribution/Results: We formally introduce and substantiate “hardware-software co-design” as the central paradigm for overcoming energy-efficiency and scalability bottlenecks, revealing a bidirectional co-evolution between AI algorithms and hardware innovations. We further prospectively examine in-memory computing and neuromorphic computing, addressing their technical pathways and practical deployment challenges. The study establishes a comprehensive landscape of AI accelerators—spanning foundational principles, evaluation methodologies, and evolutionary trends—and distills universal design guidelines for high-performance, energy-efficient AI hardware, thereby providing theoretical foundations and practical guidance for next-generation AI system architectures.

Computer architecture must evolve to accelerate modern AI workloads efficientlyHardware-software co-design is necessary for future AI computing progressTraditional architectures struggle with massive computational demands of complex AI models

The rapid deployment of machine learning across platforms from milliwatt-class TinyML devices to large language models has made energy efficiency a primary constraint for sustainable AI. Across these scales, performance and energy are increasingly limited by data movement and memory-system behavior rather than by arithmetic throughput alone. This work reviews energy efficient software hardware codesign methods spanning edge inference and training to datacenter-scale LLM serving, covering accelerator architectures (e.g., ASIC/FPGA dataflows, processing-/compute-in-memory designs) and system-level techniques (e.g., partitioning, quantization, scheduling, and runtime adaptation). We distill common design levers and trade-offs, and highlight recurring gaps including limited cross-platform generalization, large and costly co-design search spaces, and inconsistent benchmarking across workloads and deployment settings. Finally, we outline a hierarchical decomposition perspective that maps optimization strategies to computational roles and supports incremental adaptation, offering practical guidance for building energy and carbon aware ML systems.

data movementenergy efficiencyhardware-software co-design

How to Keep Pushing ML Accelerator Performance? Know Your Rooflines!

May 22, 2025
MV
Marian Verhelst
🏛️ KU Leuven | imec | ETH Zurich | Universit.a di Bologna | Princeton University | EnCharge AI

To address the performance bottlenecks of machine learning (ML) accelerators under growing model sizes and stringent energy-efficiency constraints, this paper proposes the first enhanced Roofline model deeply co-designed for ML accelerator characteristics. Our method introduces *execution paradigm boundary analysis* and *energy-constrained performance upper-bound modeling*, unifying the quantification of computational intensity, memory hierarchy, and data layout effects on both performance and energy efficiency. By integrating ML workload feature extraction, architecture-level quantitative analysis, and a hardware-algorithm co-evaluation framework, we systematically identify performance bottlenecks across mainstream accelerators for diverse operators and memory layouts. Experimental results reveal synergistic optimization pathways between memory bandwidth and computational density, and clarify several open research directions. The proposed model provides both theoretical foundations and practical guidance for energy-aware ML accelerator architecture design.

Applying roofline model to optimize execution regimesEnhancing ML accelerator performance and efficiencyUnderstanding compute-memory interactions for system efficiency

This work addresses the challenge of prolonged and error-prone manual kernel development for emerging AI accelerators, which stems from their use of specialized instruction set architectures (ISAs) and hinders cross-platform portability. To overcome this, the paper introduces the first agent-driven benchmark for kernel generation tailored to novel hardware, featuring a large language model (LLM)-based feedback optimization framework. This framework leverages function calling and iterative refinement to automatically synthesize efficient and correct low-level kernels. Evaluation across more than twenty machine learning tasks on three distinct emerging accelerators demonstrates that the approach rapidly generates high-performance kernel code—often matching or surpassing compiler-generated baselines—even for previously unseen ISAs, thereby significantly accelerating the hardware development cycle.

AI acceleratorsemerging hardwareinstruction set architecture

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

Nov 20, 2024
JF
Jared Fernandez
🏛️ Meta | Carnegie Mellon University

Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.

Assessing diminishing returns in scaling accelerators for large model trainingEvaluating parallelization strategies to minimize distributed communication overheadOptimizing hardware configuration for efficient large-scale model training

Latest Papers

What's happening recently
View more

This work addresses the challenges of efficiently deploying AI models under real-time and energy-constrained conditions, where conventional CPU/GPU platforms fall short and FPGA-based solutions remain hindered by complex hardware-software co-design and data scheduling. To overcome these limitations, the authors propose the AI FPGA Agent framework, which introduces a novel runtime software agent mechanism to enable dynamic model partitioning, hardware task scheduling for compute-intensive layers, and automated data transfer management. Integrated with a configurable quantized arithmetic acceleration core, this approach yields a low-intervention, energy-efficient, and reconfigurable FPGA inference system. Experimental results demonstrate over 10× latency reduction compared to CPU baselines and 2–3× higher energy efficiency than GPUs, while keeping classification accuracy degradation within 0.2%.

AI-FPGA integrationdeep neural network inferenceenergy-efficient AI

Hardware optimization on Android for inference of AI models

Nov 17, 2025
IG
Iulius Gherasim
🏛️ Complutense University of Madrid

To address the challenge of efficiently orchestrating heterogeneous hardware (e.g., GPU, NPU) for real-time AI inference on Android devices, this paper proposes a mobile-oriented, hardware-aware inference optimization framework. Methodologically, it integrates model quantization (INT8/FP16), operator-level hardware mapping, and coordinated accelerator scheduling to derive optimal execution configurations for YOLO (object detection) and ResNet (image classification) models. Its key contributions are: (1) the first system-level joint optimization of quantization accuracy, hardware resource utilization, and inference latency on Android; and (2) a lightweight configuration search mechanism enabling rapid cross-platform adaptation across Qualcomm, MediaTek, and Huawei NPUs. Experiments demonstrate that, with <1.2% mAP/Top-1 accuracy degradation, the framework achieves an average 58.3% reduction in end-to-end inference latency and a 2.1× improvement in energy efficiency—significantly outperforming TensorFlow Lite and ONNX Runtime mobile deployments.

Balancing accuracy preservation with inference speed accelerationEvaluating quantization and accelerator usage for mobile AIOptimizing AI model inference latency on Android hardware

This work addresses the multi-objective optimization challenge of deploying AI models on ARM Cortex-M embedded processors, balancing energy efficiency, accuracy, and resource utilization. The authors propose an automated, Pareto-optimal multi-objective benchmarking framework to systematically evaluate key performance metrics across Cortex-M0+, M4, and M7 cores. Their analysis reveals an approximately linear relationship between FLOPs and inference latency and quantifies the trade-off between energy consumption and model accuracy through Pareto front characterization. The study demonstrates that the M7 excels in short inference tasks, the M4 achieves superior energy efficiency for longer workloads, and the M0+ is best suited for lightweight applications, thereby offering clear guidance for processor selection and sustainable design in embedded AI systems.

AI benchmarkingARM Cortex processorsembedded systems

This work addresses the challenge of deploying low-latency, small-batch neural networks with fully on-chip weight storage in extreme-edge scientific computing, where traditional programmable logic falls short. The study systematically evaluates the inference performance gap between AI Engines and programmable logic, proposing spatiotemporal dataflow optimizations tailored for AI Engines. It introduces a novel metric—Latency-Adjusted Resource Equivalence (LARE)—to rigorously delineate, for the first time, the operational regimes where AI Engines outperform programmable logic. Through comprehensive architectural analysis, microbenchmarking, spatial and API-level optimizations, and integration into the hls4ml toolchain, the authors successfully deploy multiple end-to-end neural networks, demonstrating the scalability and performance advantages of AI Engines under stringent resource constraints.

AI Enginesextreme-edge computinglatency