Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文分析了AI代理在处理并发请求时的资源动态特性,并提出基于任务特性的CPU分配方法,以优化资源利用和降低延迟。
📝 Abstract
LLM-based AI agents process user requests through iterative reasoning and tool execution, often involving the invocation of remote LLM APIs with local tool containers. This execution model can make the optimization of agent serving difficult because latency, local resource demand, and container bottlenecks inter-mix across requests. However, the current agent ecosystem runs without much consideration of resource dynamics, which results in significant waste of the precious resources. This paper analyzes the resource inter-mix of AI agents for three representative tasks: retrieval-augmented question answering, web search, and software coding. To this end, we characterize the latency with respect to the resource dynamics of processing multiple requests and tasks concurrently. Our measurements show that agents have a wide range of behaviors depending on tasks, so that even the same tool can differ substantially in resource dynamics. We also find that running multiple requests concurrently exposes task-dependent bottlenecks in resource dynamics such as CPU, disk I/O, and memory. Furthermore, we uncover that faster LLM responses or more CPU cores do not always accelerate agents. Based on these observations, we demonstrate new optimization opportunities that exploit the resource dynamics of tasks: CPU-aware tool admission and task-aware CPU allocation. Our results show that the latency of CPU-sensitive agent tasks improves $\sim$5.4$\times$, and the average latency across multiple tasks is reduced $\sim$32% compared to native agents.
Problem

Research questions and friction points this paper is trying to address.

AI Agents
Resource Dynamics
Latency
Optimization
Container Bottlenecks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Resource Dynamics
CPU-aware Tool Admission
Task-aware CPU Allocation
W
Wonmi Choi
Department of Computer Science and Engineering, Korea University
M
Minuk Park
Department of Computer Science and Engineering, Korea University
Zhixiong Niu
Zhixiong Niu
Microsoft Research
DatacenterInternet
Yongqiang Xiong
Yongqiang Xiong
Microsoft Research Asia
Computer networkingOperating Systems
C
Chuck Yoo
Department of Computer Science and Engineering, Korea University
Gyeongsik Yang
Gyeongsik Yang
Korea University
Operating systemsNetwork virtualizationDatacenter networkingDistributed deep learning