🤖 AI Summary
This study addresses the reliability challenges of large language model (LLM) agents in tool invocation, particularly under constrained interaction budgets, where improper interface design often leads to misuse. For the first time, interface format is treated as an independent variable, and a controlled experiment systematically compares three contract representations—free-form documentation, JSON Schema, and schema augmented with structured diagnostic feedback—while holding semantic content constant. The evaluation employs a deterministic sandbox environment, structured validation diagnostics, and a fully crossed design across multiple models, random seeds, and budget levels. Results show that structured schemas significantly reduce syntactic misuse but fail to mitigate semantic errors. Critically, task success rates remain at zero across all conditions, revealing that semantic misjudgment and time constraints constitute the primary bottlenecks in current LLM-based tool use.
📝 Abstract
Tool use has become central to modern LLM agents, yet interface design is rarely isolated as an experimental variable. This paper studies whether schema based tool contracts and structured validation diagnostics improve reliability under strict interaction budgets. We evaluate three conditions that preserve identical tool semantics and information content: free form documentation, JSON Schema specifications, and JSON Schema with structured diagnostics.
We implement a deterministic software engineering sandbox with logs, metrics, configurations, and repository tasks, and evaluate a fully crossed pilot with one open local model, three seeds, three interface conditions, and four budgets. We report end task success, interface misuse, execution failures, semantic misuse, recovery behavior, and overhead. In this pilot, success remains zero across conditions, while schema conditions reduce interface misuse but not semantic misuse. The evidence supports a precise interpretation that interface formalization improves contract adherence, but semantic action quality and timeout sensitive tasks remain dominant bottlenecks under constrained local inference.