Agent Behavior as Code: Efficient and Robust LLM Agents with Programmatic Specifications

πŸ“… 2026-10-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenges of behavioral instability, high reasoning costs, and difficult long-context management in large language model (LLM) agents by proposing the ABCAgent framework. This method employs a hybrid neural-symbolic architecture that explicitly specifies agent behaviors through Python symbolic programs while avoiding premature variable binding. Within this framework, LLMs are dedicated to code editing and generation, thereby balancing execution flexibility with determinism. Experimental results demonstrate that ABCAgent achieves over 98% accuracy on benchmarks such as GSM-Symbolic, reduces latency by 5.2Γ—, and decreases invocation costs by 7Γ—. These findings indicate that the proposed approach significantly enhances both the robustness and operational efficiency of LLM-based agents.
πŸ“ Abstract
AI agents based on foundation models (FMs) have demonstrated strong capabilities to perform complex open-ended tasks. However, they face some common challenges in practice: (a) agent behavior can deviate drastically even for semantically similar tasks, leading to catastrophically propagated errors; (b) high cost and latency due to FM calls, repeated in full whenever a task recurs with different inputs; (c) FMs'limited context and instruction following capability confine how well agents manage the ever-growing execution context and follow complex plans. We introduce $\textbf{A}$gent $\textbf{B}$ehavior as $\textbf{C}$ode $\textbf{Agent}$ (ABCAgent), which uses a symbolic program (e.g., Python code with potential neural functions) to fully specify the agent's behavior at runtime, with a powerful FM agent editing that program for flexibility. Behavior is thus specified without premature variable binding, and its execution is deterministic. We evaluate ABCAgent on six agent benchmarks, two of which we construct to test how well a derived program generalizes to variants of the task it was written for. ABCAgent matches a model-matched neural agent on GAIA and augmented GAIA, and surpasses it where robustness and long control flows matter: 98.3% against 97.3% on GSM-Symbolic ($p = 0.001$), 71.9% against 47.4% $\mathrm{Pass}^4$ on the telecom domain of $\tau^2$-bench ($p = 0.0001$), and more records written correctly at every loop length on our control-flow-augmented WorkArena benchmark. For more parametric task families, ABCAgent is also significantly superior in efficiency. Without authoring a new program, ABCAgent solves 92.6% of GSM-Symbolic instances and 20.1% of augmented GAIA variants, which yields $5.2\times$ lower latency and $7.0\times$ lower cost on GSM-Symbolic, 19% lower cost on augmented GAIA, and $9.5\times$ lower agent latency on $\tau^2$-telecom.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
behavioral robustness
cost and latency
context limitation
instruction following
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agent Behavior as Code
Symbolic Program
Neuro-symbolic Agent
Deterministic Execution
LLM Agent Efficiency
P
Peng Qi
Uniphore
C
Chunliang Lyu
Uniphore
G
Gang Li
Uniphore
F
Fabian Chan
Uniphore
C
Cheng Chang
Uniphore
Ignacio Cases
Ignacio Cases
Postdoc at CSAIL, MIT
Computational LinguisticsDeep Reinforcement LearningNLU
W
Will Lu
Uniphore