PAI: Fast, Accurate, and Full Benchmark Performance Projection with AI

๐Ÿ“… 2026-03-18
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Traditional performance simulators are slow and costly, while existing machine learning approaches suffer from limitations in accuracy, speed, or coverage. This work proposes a hierarchical LSTM-based AI model that leverages microarchitecture-agnostic program execution feature traces to efficiently predict full-benchmark performance metrics without requiring detailed simulation or instruction-level encoding. The method achieves, for the first time, high-accuracy, end-to-end whole-program performance prediction, attaining an average IPC prediction error of only 9.35% on SPEC CPU 2017. The entire prediction process completes in just 2 minutes and 57 secondsโ€”offering accuracy comparable to state-of-the-art techniques while accelerating prediction by three orders of magnitude.

Technology Category

Machine Learning: Hardware-aware MLHumans and AI: Human-in-the-loop Machine LearningCognitive Modeling & Cognitive Systems: Agent Architectures

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
๐Ÿ“ Abstract
The exponential increase in complex IPs within modern SoCs, driven by Moore's Law, has created a pressing need for fast and accurate hardware-software power-performance analysis. Traditional performance simulators (such as cycle accurate simulators) are often too slow to simulate full benchmarks within a reasonable timeframe; require considerable effort for development, maintenance, and extensions; and are prone to errors, making pre-silicon performance projections and competitive analysis increasingly challenging. Prior attempts in addressing this challenge using machine learning fall short as they are either slow, inaccurate or unable to predict the performance of full benchmarks. To address these limitations, we present PAI, the first technique to accurately predict full benchmark performance without relying on detailed simulation or instruction-wise encoding. At the heart of PAI is a hierarchical Long Short Term Memory (LSTM)-based model that takes a trace of microarchitecture independent features from a program execution and predicts performance metrics. We present the detailed design, implementation and evaluation of PAI. Our initial experiments showed that PAI can achieve an average IPC prediction error of 9.35% for SPEC CPU 2017 benchmark suite while taking only 2 min 57 sec for the entire suite. This prediction error is comparable to prior state-of-the-art techniques while requiring 3 orders of magnitude less time.
Problem

Research questions and friction points this paper is trying to address.

performance projection
full benchmark
hardware-software co-analysis
pre-silicon validation
SoC performance modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

performance prediction
LSTM
microarchitecture-independent features
full benchmark
AI-driven simulation
๐Ÿ”Ž Similar Papers
No similar papers found.
A
Avery Johnson
Texas A&M University, College Station, TX 77840 USA
M
Mohammad Majharul Islam
Intel Corporation, Hillsboro, OR 97124 USA
R
Riad Akram
Intel Corporation, Hillsboro, OR 97124 USA
Abdullah Muzahid
Abdullah Muzahid
Professor of Computer Science and Engineering, Texas A&M University
Computer Architecture