Can LLM Agents Price Competitively? A Dynamic Multi-Attribute Auction Benchmark for Agentic Commerce

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the pricing capabilities of large language model (LLM)-driven agents in complex, dynamic markets characterized by hidden customer preferences, real-time competitor responses, and abrupt demand shifts. To this end, we introduce Bazaar, the first dynamic multi-attribute sealed-bid auction benchmark for autonomous commercial agents, which integrates closed-form customer utility functions to balance real-world market complexity with evaluability. Through multi-agent simulations grounded in dynamic game-theoretic modeling, we systematically evaluate eleven state-of-the-art LLMs and uncover a pronounced trade-off between customer acquisition and profit maximization, alongside marked disparities in responsiveness to demand shocks. Notably, even the best-performing LLM agent achieves less than one-third of the ex post optimal profit, highlighting substantial room for improvement in LLM-based autonomous business decision-making.
📝 Abstract
Agentic commerce is moving from concept to deployed infrastructure: payment networks, retailers, and AI platforms are setting the stage for agents to transact on behalf of merchants and consumers. Yet whether the LLMs behind these agents can price competently in real markets, where customer preferences are hidden, competitors adapt in real time, and demand can shift without warning, has not been systematically tested. We introduce Bazaar, a dynamic sealed-bid benchmark for multi-attribute auction under these conditions. Despite its dynamics, the benchmark is grounded in closed-form customer utilities, enabling exact evaluation. Across 11 frontier LLMs from four providers, the leading agents on customer acquisition (e.g. Gemini 3.1 Pro) are often not the leading agents on profit (e.g. Opus 4.6). The ranking shifts again under demand shocks: agents that learned fastest pre-shock are typically the slowest to revise their beliefs afterwards, while Gemini 3.1 Pro recovers fastest despite not leading on profit. However, even the strongest agent captures less than a third of hindsight-optimal profit, suggesting current LLMs are progressing in agentic commerce but leave substantial headroom.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
dynamic pricing
multi-attribute auction
agentic commerce
competitive pricing
Innovation

Methods, ideas, or system contributions that make the work stand out.

dynamic multi-attribute auction
agentic commerce
LLM agents
pricing competence
Bazaar benchmark
🔎 Similar Papers
2024-03-31arXiv.orgCitations: 12