Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the behavioral drivers behind unsafe development practices in artificial intelligence competitions, where developers often compromise safety under competitive pressure. Through a behavioral experiment simulating an AI development race, combined with risk preference assessments and a repeated-game design, the authors construct an evolutionary game-theoretic model incorporating four strategic types. The findings reveal that unsafe development is primarily driven by opponents’ behavior, fear of falling behind, and early behavioral inertia—rather than individual risk preferences. Experimental results demonstrate that falling behind significantly increases the likelihood of unsafe choices, whereas leading suppresses such behavior. The model successfully replicates these treatment effects, illustrating how the competitive structure itself can generate conditional unsafe strategies, thereby offering a novel mechanistic explanation for safety trade-offs in AI development races.
📝 Abstract
Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-conscious development. We study this using a framed behavioural experiment based on an idealised AI race, in which paired participants repeatedly chose between Safe and Unsafe development under an uncertain time horizon. Unsafe development gave faster progress and higher immediate payoffs but accumulated private risk up to a treatment-specific maximum of 10\%, 60\%, or 90\%; the race's competitive structure was held constant, and only this maximum risk varied. Neither the pre-registered comparison between risk levels nor the role of elicited risk preferences was supported by the data. Instead, exploratory analyses motivated by the task's repeated structure show that Unsafe behaviour is shaped less by risk preferences than by the evolving strategic state of the race: participants are more likely to choose Unsafe after their opponent does so, being ahead reduces Unsafe play while falling behind increases it, and first-round choices predict later behaviour. To interpret these effects we introduce a reduced evolutionary model with four strategies -- Always Safe, Always Unsafe, Conditionally Safe, and Conditionally Antisocial Safe -- which reproduces the treatment effect and shows how conditional Unsafe behaviour can be favoured by competitive race dynamics. Together, the experiment and model show that unsafe development can emerge from early behavioural momentum, opponent behaviour, and fear of falling behind, rather than from risk preferences alone, suggesting policy should focus on reducing competitive pressure and promoting cooperation in AI development rather than only individual risk.
Problem

Research questions and friction points this paper is trying to address.

AI race
unsafe development
competitive pressure
risk preferences
strategic behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI race
behavioral experiment
competitive pressure
evolutionary model
unsafe development
🔎 Similar Papers