🤖 AI Summary
This study addresses the persistent challenge that large-scale AI systems often achieve formal compliance while still generating substantive harms. It introduces a political economy framework into AI accountability research, proposing a sequential game-theoretic model to characterize strategic interactions among AI vendors, deployers, and regulators under conditions of switching costs and evidentiary dependence. The analysis reveals the conditions under which an “agency-compliant” equilibrium emerges. By integrating game theory, mechanism design, and institutional analysis, the work identifies a unique interior equilibrium alongside a corner solution that fully mitigates harm, thereby explaining why standardized evaluations frequently fail to curb ongoing adverse impacts. Furthermore, it systematically evaluates the incentive effects of institutional mechanisms—including independent auditing, model portability, incident reporting, and outcome-based joint liability—and proposes empirically testable pathways for regulatory intervention.
📝 Abstract
AI accountability at scale is an institutional problem: who can observe, verify, and change deployed systems. We develop a sequential political-economy model in which an AI vendor chooses auditability and substantive mitigation, a deployer monitors after adoption while facing switching costs, and enforcement depends on verifiable evidence. Anticipating the deployer's monitoring response, the vendor may stop at an observable procurement floor while mitigating below the social first best, producing a proxy-compliance equilibrium. We characterize the unique interior equilibrium and the corner in which harm is fully mitigated. Independent audit rights raise enforcement exposure directly; portability restores deployer leverage; incident reporting adds a regulator-visible evidence channel; and outcome-linked liability creates incentives that do not depend on vendor-controlled detection. The results explain why documentation and standardized evaluations can coexist with persistent post-deployment harms, and generate testable implications for monitoring, mitigation, and the gap between formal compliance and operational outcomes.