🤖 AI Summary
This study addresses the resource allocation imbalances and collective inefficiencies arising from prediction homogenization among AI agents sharing common models. Employing a two-route congestion game framework, we conduct multi-agent simulations comparing model families such as GPT, alongside human-AI hybrid experiments. Our findings reveal that shared predictions induce pronounced herding behavior in AI populations, leading to severe congestion, whereas human decision-making maintains equilibrium. In hybrid settings, individual burdens become highly unequal yet remain obscured by average cost metrics. This work demonstrates that evaluating AI systems must transcend mean-based indicators to account for collective dynamics and the fairness of cost distributions.
📝 Abstract
AI agents built on a few shared models increasingly act for many people. A shared forecast about others can align their choices and change how scarce capacity is allocated. We tested this feedback in a two-road congestion game. Adding one sentence warning that others might follow a routing tip made populations of 50 GPT agents crowd one road while avoiding the nearly empty alternative. Average travel time rose from 64 to 95 min, although any crowded-road agent could have saved 69 min by switching alone. The warning discouraged the very move it predicted. The pattern persisted for 100 rounds. Two other model families shifted the same way without locking onto one road. Twelve all-human groups (240 participants) stayed near balance under numerical reports or the tip and warning. In 24 mixed groups with a further 240 participants, imbalance grew with the share of agents in the registered analysis, while people increasingly took the road the agents avoided. Collective costs stayed below the allagent reference, but with 15 agents and 5 humans, agent seats averaged 80 min, compared with 44 min for human seats. Shared forecasts can thus sustain collective inefficiency among similar agents. A better group average can also hide an unequal burden. Evaluations of AI agents that share resources should test populations, treat messages as interventions and report who bears the costs.