AI Hardware Accelerators for Large Language Models: Architectures and the Memory Wall

📅 2026-08-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了使用不同硬件加速器解决大型语言模型训练和服务中的内存墙问题,通过比较计算、内存、能耗等特性,提出处理内存问题是关键。
📝 Abstract
Large language models (LLMs) place unprecedented and still-growing demands on the hardware that trains and serves them. This review surveys the full landscape of AI hardware accelerators for LLMs, including general-purpose GPUs, custom ASICs such as TPUs, Trainium, Groq, and Cerebras, reconfigurable FPGAs, processing-in-memory and near-memory architectures, and emerging neuromorphic and photonic approaches across cloud and edge deployment. Using the transformer's computational structure and roofline analysis as a common framework, we show that the decisive constraint on LLM acceleration is not arithmetic but memory: the autoregressive decode phase is bandwidth-bound, the key-value cache can rival the model weights in size, and data movement dominates energy. Comparing platforms on compute, memory, energy, programmability, and scalability, we find that no single architecture is optimal across workloads: GPUs remain the flexible default and the workhorse of training; domain-specific ASICs win at scale for stable, high-volume workloads; processing-in-memory is the most promising near-term response to the memory wall, entering systems as a heterogeneous complement; and neuromorphic and photonic computing, while promising, are not yet production-ready at frontier scale. Future progress depends on hardware-algorithm co-design and heterogeneous, memory-centric systems: for large language models, the memory system has become the computer.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Memory Wall
Hardware Accelerators
Innovation

Methods, ideas, or system contributions that make the work stand out.

memory wall
processing-in-memory
hardware-algorithm co-design