🤖 AI Summary
This work addresses the architectural overfitting caused by the limited scale of benchmarks such as SPEC, as well as the prohibitive cost of manually constructing large-scale simulation benchmarks. To overcome these challenges, we propose an agent-driven automated workflow that leverages AI agents to transform open-source code repositories end-to-end into microarchitecture simulation-ready executables, thereby replacing traditional manual curation pipelines. Experimental results demonstrate that this system can rapidly construct diverse benchmark suites comprising hundreds of applications. Compared to SPEC, the generated benchmarks deliver higher-fidelity performance evaluations and reveal novel architectural design insights.
📝 Abstract
The selection of benchmark workloads is of paramount importance in computer architecture, as it establishes the yardstick against which architectural innovations are measured and guided. Yet for decades, the SPEC benchmark suites, comprising merely tens of workloads, have been the de facto standard in academic architectural research, where they are frequently treated as a principal evaluation and optimization target. When a suite this small is relied upon so heavily, it risks architectural overfitting; as our research and prior studies demonstrate, an overly narrow focus can mislead design decisions by overvaluing certain innovations, producing cores that excel on SPEC benchmarks yet underperform on broader, realistic workloads.
To mitigate this overfitting, adopting a large, comprehensive benchmark suite is the natural solution. However, the immense engineering effort required to strip software into the clean, interference-free binary executables demanded by simulators often makes this highly impractical. In this work, we demonstrate that AI agents provide an elegant solution to this challenge. Rather than manually curating yet another static benchmark suite, we introduce an agent-driven workflow capable of autonomously transforming arbitrary open-source repositories into simulator-ready executables. This automated approach makes workload collection highly scalable, allowing us to rapidly harvest hundreds of diverse applications from public repositories into our benchmark suite. Through a comparative analysis of our agent-generated suite against SPEC, we show that it not only achieves higher-fidelity performance assessments but also uncovers novel architectural insights that traditional, static suites fail to expose.