🤖 AI Summary
This study addresses the fragmentation in AI alignment research stemming from competing conceptual frameworks, which can lead interventions to have opposing effects under different alignment perspectives. By systematically analyzing three dominant alignment paradigms, the work reveals that their fundamental disagreements arise from divergent threat models and normative orientations. Through conceptual analysis, comparison of research programs, and clarification of policy–science distinctions, the paper articulates— for the first time—the internal pluralism and tensions within alignment discourse. It proposes five recommendations to improve research practices and makes a key contribution by developing a refined conceptual framework that distinguishes idealized alignment goals from empirical proxy metrics. This framework provides a clearer terminological foundation and methodological guidance for interdisciplinary communication and technical intervention in AI alignment.
📝 Abstract
The ML literature contains many distinct concepts falling under the heading of 'AI alignment'. After noting three concepts of AI alignment in the context of their corresponding research programs, we claim that realistic interventions may promote 'AI alignment' under one conception while being actively counterproductive from the perspective of others. We suggest that tensions between alignment ideals emerge due to differences in background threat-models, alongside differences in normative orientations. In light of our analysis, researchers aiming to further the goal of 'AI alignment' should do five things. First, they should not conflate distinctions of policy and distinctions of scientific scope; second, methodological disagreements should be acknowledged explicitly; third, researchers should distinguish between 'AI alignment' as a high-level ideal and specific 'alignment proxies' used in empirical research; fourth, they should use more granular concepts to identify both the source and nature of possible AI harms/benefits; fifth, they should explicitly acknowledge the diversity of 'alignment' concepts in both empirical work and in communication with non-technical audiences.