How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过对比预写任务计划与随机策略文本,分析了代理约束在零售和航空实验中如何影响成功率、错误接受率及成本。
📝 Abstract
Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airline pilot in $τ^2$-bench. The primary comparison pairs prewritten task-specific plans (Fixed) with shuffled policy text matched in word count (Sham), isolating the contribution of guidance content. Across 265 matched cells, Fixed improves oracle-verified success by 7.17 percentage points (90\% task-clustered bootstrap interval, 1.15--13.36 points), with gains concentrated in higher-complexity tasks. A read-only terminal verifier rejects 61\% of Retail oracle-invalid episodes while withholding 17\% of correct ones, at less than one cent of additional cost per episode. Which component matters more depends on the loss assigned to erroneous acceptance: at low liability the planning gain dominates; at high liability the verifier's avoided false passes dominate---and a standalone verifier captures nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction of its cost.
Problem

Research questions and friction points this paper is trying to address.

Agent Harnesses
Planning Information
Release Control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agent Harnesses
Planning Information
Release Control
Stateful LLM Agents
Error Acceptance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yukun Zhang
Yukun Zhang
哈尔滨工业大学(深圳)
computer scienceai
K
Kemu Xu
University of Edinburgh
Y
Yishen Chen
The Chinese University of Hong Kong, Shenzhen