Schema First Tool APIs for LLM Agents: A Controlled Study of Tool Misuse, Recovery, and Budgeted Performance

📅 2026-03-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliability challenges of large language model (LLM) agents in tool invocation, particularly under constrained interaction budgets, where improper interface design often leads to misuse. For the first time, interface format is treated as an independent variable, and a controlled experiment systematically compares three contract representations—free-form documentation, JSON Schema, and schema augmented with structured diagnostic feedback—while holding semantic content constant. The evaluation employs a deterministic sandbox environment, structured validation diagnostics, and a fully crossed design across multiple models, random seeds, and budget levels. Results show that structured schemas significantly reduce syntactic misuse but fail to mitigate semantic errors. Critically, task success rates remain at zero across all conditions, revealing that semantic misjudgment and time constraints constitute the primary bottlenecks in current LLM-based tool use.

Technology Category

Planning, Routing, and Scheduling: Planning with Language ModelsMachine Learning: Large Multimodal Models (LMMs)Humans and AI: Intelligent User Interfaces

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Tool use has become central to modern LLM agents, yet interface design is rarely isolated as an experimental variable. This paper studies whether schema based tool contracts and structured validation diagnostics improve reliability under strict interaction budgets. We evaluate three conditions that preserve identical tool semantics and information content: free form documentation, JSON Schema specifications, and JSON Schema with structured diagnostics. We implement a deterministic software engineering sandbox with logs, metrics, configurations, and repository tasks, and evaluate a fully crossed pilot with one open local model, three seeds, three interface conditions, and four budgets. We report end task success, interface misuse, execution failures, semantic misuse, recovery behavior, and overhead. In this pilot, success remains zero across conditions, while schema conditions reduce interface misuse but not semantic misuse. The evidence supports a precise interpretation that interface formalization improves contract adherence, but semantic action quality and timeout sensitive tasks remain dominant bottlenecks under constrained local inference.
Problem

Research questions and friction points this paper is trying to address.

tool misuse
LLM agents
interface design
interaction budgets
schema contracts
Innovation

Methods, ideas, or system contributions that make the work stand out.

JSON Schema
tool misuse
structured diagnostics
LLM agents
interaction budget
A
Akshey Sigdel
Independent Researcher
R
Rista Baral
Independent Researcher