🤖 AI Summary
The computational demands of AI models—particularly deep neural networks (DNNs)—continue to outpace the capabilities of conventional architectures. Method: This paper systematically analyzes design principles and performance trade-offs across mainstream AI accelerators—including GPUs, ASICs, and FPGAs—employing architectural analysis, fine-grained performance modeling, and multi-dimensional benchmarking to investigate key techniques: dataflow optimization, memory hierarchy restructuring, sparsity exploitation, and low-precision quantization. Contribution/Results: We formally introduce and substantiate “hardware-software co-design” as the central paradigm for overcoming energy-efficiency and scalability bottlenecks, revealing a bidirectional co-evolution between AI algorithms and hardware innovations. We further prospectively examine in-memory computing and neuromorphic computing, addressing their technical pathways and practical deployment challenges. The study establishes a comprehensive landscape of AI accelerators—spanning foundational principles, evaluation methodologies, and evolutionary trends—and distills universal design guidelines for high-performance, energy-efficient AI hardware, thereby providing theoretical foundations and practical guidance for next-generation AI system architectures.
📝 Abstract
The remarkable progress in Artificial Intelligence (AI) is foundation-ally linked to a concurrent revolution in computer architecture. As AI models, particularly Deep Neural Networks (DNNs), have grown in complexity, their massive computational demands have pushed traditional architectures to their limits. This paper provides a structured review of this co-evolution, analyzing the architectural landscape designed to accelerate modern AI workloads. We explore the dominant architectural paradigms Graphics Processing Units (GPUs), Appli-cation-Specific Integrated Circuits (ASICs), and Field-Programmable Gate Ar-rays (FPGAs) by breaking down their design philosophies, key features, and per-formance trade-offs. The core principles essential for performance and energy efficiency, including dataflow optimization, advanced memory hierarchies, spar-sity, and quantization, are analyzed. Furthermore, this paper looks ahead to emerging technologies such as Processing-in-Memory (PIM) and neuromorphic computing, which may redefine future computation. By synthesizing architec-tural principles with quantitative performance data from industry-standard benchmarks, this survey presents a comprehensive picture of the AI accelerator landscape. We conclude that AI and computer architecture are in a symbiotic relationship, where hardware-software co-design is no longer an optimization but a necessity for future progress in computing.