🤖 AI Summary
This study investigates how descriptive text in LLM tool registries interferes with AI agents’ tool selection decisions through rhetoric, positioning, and structure. Employing a preregistered experimental design via the OpenAI API, we construct controlled tool-list pairs containing marketing language, canonical fields, and varied orderings to quantify, for the first time, the causal effects of “stacked praise” and primacy bias on model invocation preferences. Our findings reveal that stacked praise and primacy effects increase selection rates by approximately 43% and 72%, respectively, while sales copy significantly undermines the facilitative role of structured fields. Accordingly, we propose defensive strategies involving the concealment of promotional language and randomized ordering, providing an empirical foundation for information design in AI agent ecosystems.
📝 Abstract
AI agents often pick tools from registries, where each tool's provider writes its description. We ask whether sales language in those descriptions changes which tool an agent picks. We built pairs of listings differing in one controlled way (added praise, a verifiable specification, or list order) and asked two OpenAI models to call one tool. In a preregistered study, stacked praise (four kinds combined) raised a tool's pick rate by about 43 percentage points, matching or beating a verifiable specification. Praise also pulled some picks toward tools that could not do the task, but rarely toward tools asking for unneeded data access. With identical listings, the first-listed tool was picked about 72 points more often. On tasks with numeric limits, structured fields helped agents pick the capable tool; adding the provider's sales text beside the fields reduced or erased that gain. Registries could list limits as fields, hide sales text from agents, and randomize order. Stacked praise, but no single kind, replicated on held-out domains. Results are provisional until blind phrase ratings are complete, and cover two small models.