🤖 AI Summary
Current discussions on AI existential risk suffer from conceptual ambiguity due to the absence of a precise definition of “control.” This work addresses this gap by proposing, for the first time in the AI context, an operational definition of control as the capacity to set and achieve objectives. Integrating core concepts from cybernetics, management control theory, and control theory—including control loops, requisite variety, and goal alignment—the study develops a systematic analytical framework. This framework clarifies the conditions under which humans can maintain control, elucidates mechanisms through which AI systems may lead to loss of control, and demonstrates that even AI systems far below superintelligence can induce varying degrees of失控. The research not only establishes a theoretical foundation for understanding AI-induced loss of control but also offers actionable recommendations for safeguarding human oversight.
📝 Abstract
At present, loss of control risks have gained much prominence in public discussion, particularly in relation to AI, with extensive discourse present among academics, frontier labs, and even governments. However, in the existing literature, the concept seems to rest on surprisingly weak foundations, where even those that discuss loss of control extensively do not first establish what control is and what exactly is being lost. Our paper aims to address these gaps. We establish a working definition of control by anchoring it to the"setting and getting of goals". Then, we discuss various aspects of control, built on foundational concepts from related fields like cybernetics, management control, and control theory. This includes who (or what) can be in control, and the things they require to be in control, such as the ability to set goals, having a functional control loop, having requisite variety, and having sufficient goal alignment. Once a framework for control is established, we then discuss how control can be lost, how AIs can contribute to such loss of control, and offer relevant recommendations for how one can maintain control. One interesting consequence of our work is that humanity, as individuals and as groups, can lose varying degrees of control as a result of AI behaviour that is far below the level of superintelligence; the potential for loss of control scenarios (as we define them) already exist, and have existed for a long time.