🤖 AI Summary
This study addresses the rapid deployment of artificial intelligence in the Global South, where evaluation and governance mechanisms lag significantly behind, marked by a critical absence of independent auditing and accountability structures. Drawing on a decade of empirical research across Latin America, Sub-Saharan Africa, and the Asia-Pacific region—integrating second- and third-party algorithmic audits, responsible AI frameworks, prevalence-based performance analysis, and regional landscape surveys—the work demonstrates that the root cause of inadequate assessment is not technical incapacity but systemic underfunding. It identifies four recurrent issues: proxy objectives supplanting true effectiveness, performance claims failing under prevalence scrutiny, model failure on populations unseen during training, and persistent structural bias despite removal of sensitive attributes. In response, the study proposes five locally grounded evaluation and governance strategies and advocates, for the first time, a novel governance pathway mandating independent evaluations driven by development and philanthropic funders.
📝 Abstract
Artificial intelligence is being deployed across the Global South at a pace matching or exceeding the Global North, yet AI governance has not kept pace, and the gap is far wider in the South. Drawing on a decade of AI audit practice across Latin America, Sub-Saharan Africa, and Asia Pacific (the only fully published second-party audit of a deployed system in the region, Robot Laura in Brazil; two completed but unreleased national audits, of a child-welfare risk model and a public-employment matching algorithm; thirteen Responsible AI Assessments; and a regional landscape analysis), this paper documents patterns in the sparse field of Global South AI evaluation and explains why it so rarely occurs. We count fewer than twenty published second- and third-party audits of deployed systems across the region over the past decade, against hundreds of documented public-sector algorithms and multibillion-dollar national AI investments. We find four cross-cutting patterns: proxy targets that substitute predictability for validity, performance claims that collapse under prevalence analysis, populations scored by models that never saw them in training, and structural bias that persists even after protected attributes are removed. We then draw five lessons for evaluation, regulation, and funding that emerge outside the regulatory, linguistic, and data conditions of the North. We argue the gap is not, at root, a capacity problem but a funding problem: capacity follows funded demand, and no actor is currently required, or funded, to hold deployed systems to account. The actor best placed to close it is the small number of development and philanthropic funders behind most consequential AI in the region, whose funding conditions can require independent evaluation where no regulator yet does. We close with recommendations for funders, governments, and the design of evaluation itself.