A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference

📅 2025-11-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the enterprise demand for low-latency, energy-efficient LLM inference in AI agent applications, this work proposes an end-to-end vertically integrated cloud inference system. Built on a cluster of 288 NorthPole neural inference accelerators, the system integrates 4-bit integer quantization, distributed memory bandwidth optimization, high-performance runtime scheduling, and containerized inference pipelines—enabling flexible deployment across model scales and variable context lengths. Deployed across 18 × 2U servers (30 kW total power), it delivers 3.7 PB/s aggregate memory bandwidth and 115 peta-ops peak compute, concurrently serving 28 users with per-user token generation latency as low as 2.8 ms. This is the first full-stack, hardware–software co-optimized inference system leveraging a large-scale neural accelerator cluster. It significantly improves energy efficiency and service density, establishing a scalable architectural paradigm for datacenter-grade LLM inference.

Technology Category

Machine Learning: Hardware-aware MLNatural Language Processing: (Large) Language ModelsCognitive Modeling & Cognitive Systems: Agent Architectures

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
A vertically integrated, end-to-end, research prototype system combines 288 NorthPole neural inference accelerator cards, offline training algorithms, a high-performance runtime stack, and a containerized inference pipeline to deliver a scalable and efficient cloud inference service. The system delivers 115 peta-ops at 4-bit integer precision and 3.7 PB/s of memory bandwidth across 18 2U servers, while consuming only 30 kW of power and weighing 730 kg in a 0.67 m^2 42U rack footprint. The system can run 3 simultaneous instances of the 8-billion-parameter open-source IBM Granite-3.3-8b-instruct model at 2,048 context length with 28 simultaneous users and a per-user inter-token latency of 2.8 ms. The system is scalable, modular, and reconfigurable, supporting various model sizes and context lengths, and is ideal for deploying agentic workflows for enterprise AI applications in existing data center (cloud, on-prem) environments. For example, the system can support 18 instances of a 3-billion-parameter model or a single instance of a 70-billion-parameter model.
Problem

Research questions and friction points this paper is trying to address.

Scalable system for low-latency LLM inference
Energy-efficient neural network acceleration architecture
Supporting multiple model sizes in data centers
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vertically integrated neural inference accelerator system
Scalable modular design for multiple model configurations
Energy-efficient architecture with high computational throughput
IBM Research
M
Michael V. DeBole
IBM Research
Rathinakumar Appuswamy
Rathinakumar Appuswamy
Senior Research Scientist, IBM Almaden Research Center
Artificial IntelligenceInformation TheoryMachine learning
Neil McGlohon
Neil McGlohon
IBM Research
B
Brian Taba
IBM Research
S
Steven K. Esser
IBM Research
F
Filipp Akopyan
IBM Research
J
John V. Arthur
IBM Research
Arnon Amir
Arnon Amir
IBM Research
A
Alexander Andreopoulos
IBM Research
P
Peter J. Carlson
IBM Research
Andrew S. Cassidy
Andrew S. Cassidy
Principal Research Scientist, IBM Research
Neural ComputationComputer Architecture
Pallab Datta
Pallab Datta
IBM Research
M
Myron D. Flickner
IBM Research
R
Rajamohan Gandhasri
IBM Research
G
Guillaume J. Garreau
IBM Research
M
Megumi Ito
IBM Research
J
Jennifer L. Klamo
IBM Research
J
Jeffrey A. Kusnitz
IBM Research
N
Nathaniel J. McClatchey
IBM Research
J
Jeffrey L. McKinstry
IBM Research
T
Tapan K. Nayak
IBM Research
C
Carlos Ortega Otero
IBM Research
H
Hartmut Penner
IBM Research
W
William P. Risk
IBM Research
J
Jun Sawada
IBM Research