🤖 AI Summary
This study addresses the stochastic failure of AI agents under shared or structured inputs, which induces collective decision biases and auditing vulnerabilities. We present the first quantitative analysis of how shared inputs compromise multi-agent randomness. Through behavioral testing, cross-model comparisons, identifier manipulation, and correlation prediction, we reveal an intrinsic mechanism whereby reasoning models depend on identifier thresholds. Our findings demonstrate that explicit instructions fail to eliminate latent correlations, causing severe participation biases in frontier models such as GPT-6 and thereby jeopardizing fairness in resource allocation. Consequently, this work establishes a core evaluation framework that integrates correlation and predictability as essential metrics for assessing AI agents, offering critical insights for the reliable deployment of multi-agent systems in high-stakes environments.
📝 Abstract
Random selection is widely used in resource allocation and auditing, making reliable implementation essential for AI-agent systems. Behavioural tests across six reasoning models uncovered threshold and divisibility rules used in identifier-based choices. For threshold-following GPT-6 Sol and Gemini 3.8 Flash, single-agent measurements prospectively predicted correlated participation under shared identifiers and biased participation under distinct identifiers with common timestamp bits. Changing dates, formats and identifier labels revealed when these predictions held. Explicit instructions to randomize independently reduced but did not eliminate shared-input correlation. To test implications for oversight, we asked four models to select customer requests randomly for human review. GPT-6 Sol approached the target rate while selecting predictably from identifiers; the others rarely selected requests. All four closely followed supplied random draws. These findings expose collective and audit vulnerabilities that selection rates alone miss, making input-dependent bias, correlation and predictability central targets for agent evaluation.