Towards Efficient HPC Systems for Agents: Challenges and Opportunities

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scheduling pressure and resource waste in high-performance computing (HPC) systems caused by the high-frequency, fine-grained workloads of coding agents. We propose a system co-design paradigm that treats agents as first-class citizens. By providing the first quantitative characterization of agent-induced HPC resource consumption, this work establishes the concept of facility-aware tenants. It integrates system measurement, agent memory management, security policy enforcement, and I/O optimization to restructure compute, storage, and security architectures. Our analysis reveals emerging challenges, including agent-driven control-plane pressure, motivating a paradigm shift from restricting agents toward co-designing infrastructure around them. Ultimately, this research provides both a theoretical framework and a technical roadmap for building efficient, secure, agent-oriented HPC infrastructures.
📝 Abstract
Coding agents have become real users of high-performance computing (HPC) systems, yet today's HPC abstractions, interfaces, and policies remain designed for human-driven workflows. In our measurement, users running coding agents are only 19.5% of the observed population, but account for 55.8% of job submissions, 29.1% of CPU core-hours, and 42.7% of GPU-hours. Agents are not simply faster humans. They issue commands at 20.8x the human rate, decompose work into fine-grained explore-modify-execute loops, and pursue open-ended goals through trial-and-error campaigns that continue through nights and weekends. These behaviors strain abstractions built for human timescales, surfacing as control-plane pressure on the scheduler, metadata-intensive I/O on bandwidth-provisioned filesystems, repeated rediscovery of what earlier sessions already learned, and new prompt-injection and policy-enforcement surfaces. Neither banning agents nor treating them as ordinary users is sustainable. We instead argue for co-design, that facilities should treat agents as first-class principals where agents become facility-aware tenants. We outline the resulting challenges and opportunities in compute, storage, agent memory, and safety.
Problem

Research questions and friction points this paper is trying to address.

High-Performance Computing (HPC)
Coding Agents
Resource Scheduling
System Co-design
AI Safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

High-Performance Computing
Coding Agents
Co-design
First-class Principals
Agent Memory
Y
Yunjia Zheng
Harvard University
B
Bintang Dwi Marthen
Harvard University
Z
Zachary Pan
Harvard University
Minghao Li
Minghao Li
Beihang University
Natural Language Processing
R
Raminder Singh
Harvard FAS Research Computing
M
Manasvita Joshi
Harvard FAS Research Computing
Minlan Yu
Minlan Yu
Harvard University
NetworkingSystemsCloud Computing
J
Juncheng Yang
Harvard University