Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of detecting price fraud concealed under legitimate identities during autonomous AI agent payments, which poses significant commercial risks. To investigate this vulnerability, the authors propose a five-layer observation model demonstrating that conventional security scanners are entirely ineffective against such fraudulent activities. Leveraging production data, they construct a benchmark dataset and develop Gordonguard, a technical framework integrating offline auditing with online inline detection. The primary contributions include establishing a taxonomy of commercial fraud and an open-source detection framework that achieves a 6.5% false positive rate. Furthermore, the work quantitatively reveals a critical deployment constraint: the operational costs associated with manual review substantially exceed the value of microtransactions, highlighting practical limitations in real-world implementation.
📝 Abstract
AI agents now hold spend authority and settle payments without per-action human confirmation. The resulting loss is often not a security failure: a counterparty with the correct domain, the correct settlement address and a genuinely delivered service can charge more than it should, and no check keyed on identity will see it. We present three artefacts for measuring and reducing that loss. First, a taxonomy of agentic commerce fraud that separates five observation levels (agent reasoning, wire, settlement rail, counterparty, principal) from the request-level and history-level evidence available at each, and records which levels can observe which attacks. Second, Agentic Commerce Bench (ACB), a benchmark of twenty fraud classes generated from production aggregates, 1,647 catalogued service operations and 1,068 settlements, of which six involve a counterparty that is exactly who it claims to be. Third, gordonguard, an open-source detector stack and offline harness with which an operator can audit an agent configuration, replay hostile counterparties without an account, and run the same detectors inline. Calibrating to a stated false-positive budget on clean training traffic gives a 6.5% clean flag rate, replicated across three independent generations, and leaves eight of twenty classes no better than chance. On the four classes a reasoning layer can observe, a widely used agent security scanner run over its jailbreak-detection panel scores zero on all four, while correctly scoring 1.0 on a jailbreak supplied as a control. A measured median payment of $0.007 places a hard constraint on deployment: one human review costs 143 times the value of the payment it examines.
Problem

Research questions and friction points this paper is trying to address.

Agentic Commerce
Fraud Detection
AI Agents
Autonomous Payment
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Commerce Fraud
Fraud Detection Benchmark
Taxonomy
Open-source Detector Stack
Autonomous AI Agents
🔎 Similar Papers