๐ค AI Summary
This study addresses a critical blind spot in existing price-level audit methodologies: they fail to detect a class of covert collusion wherein agents coordinate by manipulating the joint distribution of unexplained bidding components while preserving individual marginal bid distributions consistent with competitive equilibrium, thereby rendering conventional single-agent detection methods entirely ineffective. The paper provides the first theoretical proof that such auditing frameworks are inherently blind to this form of collusion by construction and proposes a novel regulatory paradigm centered on identity consolidation and behavioral clustering. Integrating probabilistic coupling theory, bidding experiments with language models, empirical analysis of Ethereum block auction data, and Herfindahl index calculations, the research reveals significant residual correlations among 20 prominent language modelsโcorrelations that can be effectively mitigated by adjusting sampling temperature. In real-world on-chain data, legitimate multi-identity strategies become indistinguishable from collusion, yet identity consolidation increases market concentration by 247.5%โ324.5%.
๐ Abstract
Empirical work on algorithmic collusion asks one question of the data: are prices supracompetitive? We show this can be answered "no" by a conspiracy that is nonetheless profitable. Consider bidding agents that couple only through the joint distribution of their unexplained bid components, leaving every agent's own bid law exactly at the competitive law. Any test whose input is a single agent's price or bid history then has power exactly equal to its false-positive rate, for every coupling strength up to comonotonicity. The published detection methodology is therefore blind to this conduct by construction rather than underpowered, and no sample size repairs it.
Three empirical results follow. First, the mechanism appears in real language-model agents: twenty models from nineteen independent developers, three deployment prompts each, show residual correlation of $+0.053$ between two deployments of one model against $+0.0001$ across models, with a 95% interval clustered by developer of $[0.030, 0.078]$, under an auditor that sees every order feature and is fitted out of sample. Second, the coupling falls monotonically as sampling temperature rises ($p=0.002$), turning a deployment parameter into a candidate mitigation. Third, on 24 days of Ethereum block-building auction data covering 77,684 bids from 39 bidders, the honest population of bidder pairs is itself so dependent that a screen held at a 5% false-positive rate must sit above a floor of $+0.50$ to $+0.81$, which is 20 to 32 times the family-wise sampling threshold and does not fall as the audit window grows. Since lawful multi-identity operation and conspiracy are behaviourally indistinguishable here, the tractable regulatory target is not detection but counting: resolving 40 bidding identities into 23 operators raises the Herfindahl index by 247.5%, and adding behavioural clusters from public bid streams reaches 324.5%.