International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty

📅 2025-03-18
📈 Citations: 0
Influential: 0
📄 PDF

career value

200K/year
🤖 AI Summary
This paper addresses existential and marginalization risks arising from the malicious use or systemic failure of general-purpose artificial intelligence (GPAI). Method: It proposes the first “conditional” international AI safety treaty framework, triggered by computationally defined thresholds. The framework integrates interdisciplinary policy modeling, quantitative risk-threshold calibration, multilateral auditing design, and incentive-compatible mechanisms, establishing a dynamic governance system led by an International AI Safety Institute (AISI) endowed with authority to suspend high-risk models and conduct multidimensional compliance oversight. Contribution/Results: Its core innovation is a scientifically grounded, adaptive treaty paradigm—uniquely synthesizing compute-based thresholds, collaborative auditing, secure deployment protocols, and institutionalized verification processes. The work delivers an operationally viable treaty draft specifying five governance workflows, validated for feasibility by domain experts in AI safety, and positioned as a critical policy interface for global governance initiatives including those of the G7 and the United Nations.

Technology Category

Application Category

📝 Abstract
The malicious use or malfunction of advanced general-purpose AI (GPAI) poses risks that, according to leading experts, could lead to the 'marginalisation or extinction of humanity.' To address these risks, there are an increasing number of proposals for international agreements on AI safety. In this paper, we review recent (2023-) proposals, identifying areas of consensus and disagreement, and drawing on related literature to assess their feasibility. We focus our discussion on risk thresholds, regulations, types of international agreement and five related processes: building scientific consensus, standardisation, auditing, verification and incentivisation. Based on this review, we propose a treaty establishing a compute threshold above which development requires rigorous oversight. This treaty would mandate complementary audits of models, information security and governance practices, overseen by an international network of AI Safety Institutes (AISIs) with authority to pause development if risks are unacceptable. Our approach combines immediately implementable measures with a flexible structure that can adapt to ongoing research.
Problem

Research questions and friction points this paper is trying to address.

Addressing risks of advanced AI causing human marginalization or extinction
Evaluating feasibility of international AI safety agreements and regulations
Proposing a treaty for compute thresholds and oversight mechanisms
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proposes compute threshold for AI oversight
Mandates audits by international AISIs network
Combines flexible structure with research adaptation