Harnessing LLMs as Agents: What Does It Cost?

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that existing LLM agent evaluations fail to quantify the computational overhead of framework mechanisms such as context and memory. It proposes the Language Agent Machine (LAM) abstraction, which explicitly models resource consumption under a fixed semantic model and establishes a formal theory of execution cost. Methodologically, this work pioneers the adaptation of classical I/O lower bounds to context traffic, revealing asymptotic separations in memory access. It further derives tight sample complexity bounds for recomputation versus verification, along with optimal checkpointing rules. Empirically, the results validate predictions regarding communication and reliability trade-offs, identifying optimal checkpoint intervals and invocation granularities for GPT-6 Astra. Ultimately, this research provides a rigorous theoretical foundation for principled agent system design.
📝 Abstract
Language-model agents increasingly rely on harnesses that manage bounded context, persistent memory, tools, verification, and repeated execution, yet existing notions of model capability do not quantify the computational resources these mechanisms consume. We introduce the Language Model Agent Machine (LAM), a resource-bounded abstraction that fixes the underlying semantic model while explicitly charging harness-level resources. We establish four classes of results. Communication: LAM execution is instancewise equivalent to red--blue pebbling under simultaneous call--transfer budgets, transferring classical I/O lower bounds to context--memory traffic. Access: memory interfaces induce asymptotic separations, including a $Θ(n)$ gap between random and non-speculative sequential access on pointer chasing. Recomputation: bit-reversal DAGs require $Θ(n^2/(C+S)+n)$ model calls with context capacity $C$ and persistent-memory capacity $S$, quantifying when stored intermediate state avoids repeated semantic computation. Reliability: we derive tight stage-local sampling bounds, exact imperfect-verification costs, and a Young--Daly-type checkpoint law with a closed-form optimal verification interval. Controlled and held-out experiments on GPT-6 Astra test communication and reliability predictions, including checkpoint optima, policy selection under programmatic checking, and tradeoffs among call granularity, logical input traffic, and reliability on chained MATH tasks. Together, these results provide a resource theory for the computational cost of language-model agent harnesses.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
computational cost
resource-bounded abstraction
agent harness
resource theory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Language Model Agent Machine
resource-bounded abstraction
red-blue pebbling
checkpoint law
resource theory
🔎 Similar Papers