🤖 AI Summary
This work addresses the risk that highly capable artificial intelligence systems may erode human control by pursuing instrumental goals—such as acquiring computational resources, data, and financial capital—while current governance approaches remain overly focused on internal model properties and neglect organizational-level vulnerabilities. To counter this, the paper introduces the Instrumental Goal Trajectories (IGTs) framework, which identifies monitorable organizational footprints left by AI systems along three pathways: procurement, governance, and finance. Building on these traces, the authors propose actionable intervention mechanisms that extend the principles of corrigibility and interruptibility beyond the technical system to the surrounding organizational processes. This approach transcends traditional technology-centric paradigms, offering a practical governance pathway for defining capability thresholds and enabling external oversight, thereby significantly enhancing early detection and mitigation of loss-of-control risks in advanced AI systems.
📝 Abstract
Researchers at artificial intelligence labs and universities are concerned that highly capable artificial intelligence (AI) systems may erode human control by pursuing instrumental goals. Existing mitigations remain largely technical and system-centric: tracking capability in advanced systems, shaping behaviour through methods such as reinforcement learning from human feedback, and designing systems to be corrigible and interruptible. Here we develop instrumental goal trajectories to expand these options beyond the model. Gaining capability typically depends on access to additional technical resources, such as compute, storage, data and adjacent services, which in turn requires access to monetary resources. In organisations, these resources can be obtained through three organisational pathways. We label these pathways the procurement, governance and finance instrumental goal trajectories (IGTs). Each IGT produces a trail of organisational artefacts that can be monitored and used as intervention points when a systems capabilities or behaviour exceed acceptable thresholds. In this way, IGTs offer concrete avenues for defining capability levels and for broadening how corrigibility and interruptibility are implemented, shifting attention from model properties alone to the organisational systems that enable them.