Action-Space Shaping for LLM Agents: Measuring and Mitigating Tool-Schema Bias

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses "schema bias," a phenomenon wherein large language model (LLM) agents exhibit significant performance disparities across functionally equivalent yet syntactically distinct tool schemas. We formally define this phenomenon and propose an executable transformation framework based on nine operators that rewrites tool schemas while preserving action semantics, enabling quantitative analysis of such bias. Furthermore, we introduce a few-shot probing method to reliably predict schema difficulty and construct a multi-model evaluation benchmark. Experimental results demonstrate that even state-of-the-art models remain severely affected by schema bias, with success rates fluctuating substantially across schema variants. Notably, existing training-based mitigation strategies prove effective only for previously seen variants, exhibiting limited generalization capability.
📝 Abstract
Large Language Models (LLMs) have shown strong performance on tool-use agentic tasks when given a fixed tool schema. Yet a tool schema is not the action space of an agent; it is merely one interface representation of it. The same executable action can be exposed through many different, functionally equivalent tool definitions, and an agent that has truly learned a task should behave consistently across them. We show that current agents often do not, a phenomenon we term schema bias. To study this systematically, we introduce an executable transformation framework that rewrites a native tool schema using nine operators, including merging and splitting tools, altering how a single tool is expressed, and distributing one action across several dependent calls. The tasks, executable actions, and reachable states remain fixed, so any change in success is attributable to the interface alone. Evaluating eleven LLMs, including two closed models, on up to 32 schema variants, we ask how large schema bias is, how it manifests, whether the difficulty of a schema variant can be predicted without a full evaluation, and whether training removes it. We find that schema bias is substantial even for the newest models: success rates range from complete failure to 97% depending solely on the schema. To reliably estimate schema difficulty, it requires running a small sample of the target queries. Training repairs a schema variant only when that variant appears in the training data.
Problem

Research questions and friction points this paper is trying to address.

Schema Bias
LLM Agents
Tool Use
Action Space
Interface Representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Schema Bias
Action-Space Shaping
Executable Transformation Framework
Tool-Use Agents
LLM Agents
🔎 Similar Papers