Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce

๐Ÿ“… 2026-08-03
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge of enabling auditable and verifiable interactions between buyer and merchant AI agents in natural languageโ€“driven โ€œvibe commerce,โ€ while preserving their private objectives and permissions. It proposes Agentic Commerce World (ACWorld), the first auditable environment supporting bidirectional autonomous agent interaction. ACWorld employs the Vibe Commerce Protocol to pre-validate agent actions and integrates large-scale product catalog search with structured task trajectory logging, thereby achieving process-level verifiability and fine-grained capability evaluation. The authors release a benchmark comprising 200 capability tasks and 60 large-catalog search tasks spanning 785,022 products. Evaluations across ten models yield average scores of 65.9%โ€“85.6% and 56.1%โ€“91.4% on the two task types, respectively, highlighting the limitations of relying solely on final-state assessments.
๐Ÿ“ Abstract
In vibe coding, people describe software in natural language and delegate implementation to AI agents. By analogy, vibe commerce allows people to express buying or selling goals in natural language and delegate the corresponding tasks to agents. Commerce, however, requires independently controlled Buyer and Merchant agents to interact in a shared market while preserving their private objectives and distinct authority. We introduce Agentic Commerce World (ACWorld), an environment for evaluating such agents across ongoing transactions. Through its Vibe Commerce Protocol (VCP), ACWorld validates agent actions before updating shared transaction state and records the resulting interactions, making agent behavior auditable and evaluation reproducible. The ACWorld Benchmark contains a 200-task capability-coverage track and a 60-task large-catalog track that searches 785,022 transactable listings. Across ten models, mean scores range from 65.9% to 85.6% and from 56.1% to 91.4%, respectively. Our analysis shows that process-level evidence is necessary: final state alone can miss evaluated errors, incomplete trajectories still retain useful process signals, and large-catalog tasks expose bottlenecks across stages.
Problem

Research questions and friction points this paper is trying to address.

vibe commerce
AI agents
auditable transactions
shared market
private objectives
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Commerce
Vibe Commerce
Auditable AI Agents
Transaction Protocol
Process-level Evaluation
๐Ÿ”Ž Similar Papers
2024-02-04IEEE Communications Surveys & TutorialsCitations: 11