The Default Trap: Rethinking Plan Evaluation in Tool-Using LLM Agents

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical bias in existing LLM agent evaluations, where low responsiveness is frequently misidentified as strong priority control, leading to distorted plan assessments. To resolve this, the work introduces the concept of the "default trap" to distinguish genuine priority-driven responses from presentation-dependent default selections. Through paired plan comparisons, large-scale decision window analyses, and component-level controlled experiments across Retail, Airline, and AgentDojo scenarios, the authors systematically dissect the planning behaviors of tool-calling agents. The findings quantitatively demonstrate that reversing option order exerts a substantial influence on default goal selection, ranging from 63.3% to 98.3%. Furthermore, this research establishes a multidimensional joint evaluation framework encompassing priority, default behavior, and task success rate, effectively overcoming the limitations inherent in traditional single-metric assessment paradigms.
📝 Abstract
An executor can respond strongly to a change in a supplied plan's priority while showing a small change in the same information-selection probability when a default-aligned whole plan is removed. We call the risk of interpreting the latter as weak responsiveness to alternative priorities the default trap. We compare paired plans that prioritize different information targets with a shared no-plan reference. An accounting identity relates these distinct behavioral contrasts. Across 3,200 decision windows on 160 selected Retail, Airline, and AgentDojo tasks, switching priorities strongly redirects two models' choices, while the two plan-versus-default contrasts differ. In 2,160 additional windows, reversing account-list order shifts default target selection by 63.3-98.3 percentage points; priority-switching effects remain 96.7-100.0 points in either order. A separate 3,240-window component study finds strong control under single priority sentences, with effects of additional text varying by group and direction. Finally, 1,080 full-task episodes yield observed success differences of -19.4 to +8.3 points relative to no plan. All Retail and Airline success intervals include zero; AgentDojo results describe four fixed application worlds. These findings support joint reporting of priority responsiveness, presentation-dependent defaults, and task success and cost.
Problem

Research questions and friction points this paper is trying to address.

Tool-Using LLM Agents
Plan Evaluation
Default Trap
Priority Responsiveness
Decision Bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Default Trap
Plan Evaluation
Tool-Using LLM Agents
Priority Responsiveness
Presentation-dependent Defaults
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Xueqi Li
Xueqi Li
Shenzhen University
J
Jingjie Ning
Carnegie Mellon University
Y
Yibo Kong
Carnegie Mellon University