Score
Design and implement mechanisms that detect when external interventions or guidance should be stopped, including termination detectors, adaptive cutoff rules, and intervention schedulers that monitor agent competence and performance; analyze termination timing and outcomes to ensure disabling guidance preserves exploration benefits and yields a smooth handoff to autonomous control.
This work addresses the risk that highly capable artificial intelligence systems may erode human control by pursuing instrumental goals—such as acquiring computational resources, data, and financial capital—while current governance approaches remain overly focused on internal model properties and neglect organizational-level vulnerabilities. To counter this, the paper introduces the Instrumental Goal Trajectories (IGTs) framework, which identifies monitorable organizational footprints left by AI systems along three pathways: procurement, governance, and finance. Building on these traces, the authors propose actionable intervention mechanisms that extend the principles of corrigibility and interruptibility beyond the technical system to the surrounding organizational processes. This approach transcends traditional technology-centric paradigms, offering a practical governance pathway for defining capability thresholds and enabling external oversight, thereby significantly enhancing early detection and mitigation of loss-of-control risks in advanced AI systems.
Existing adaptive agent-based regulatory simulations lack mechanisms to systematically integrate diagnostic feedback into policy controllers, resulting in delayed and opaque policy adjustments. This work proposes a lightweight machine-guided policy revision layer that represents policies as defeasible rules and combines symbolic control, defeasible logic, and policy prioritization to operationalize contestability at the controller level, thereby endowing policy decisions with explainability, contestability, and dynamic revisability. In an emissions regulation agent-based model, the approach significantly reduces problem recurrence under scenarios where the VPVA mechanism fails due to excessive conservatism, while effectively maintaining key performance indicators such as violation rates, overshoot, and volatility.
Rapidly evolving AI systems pose systemic risks—including loss of control, misuse, geopolitical instability, and concentration of power—necessitating coordinated global governance. Method: This paper proposes a multilaterally coordinated AI activity pause mechanism, introducing four core technical modules: interruptible AI architecture design, deployment-level control protocols, full-lifecycle model tracking and certification, and a decentralized monitoring network. Contribution/Results: Moving beyond policy-centric governance, the framework establishes a verifiable, executable, and fault-tolerant technical foundation resilient to single-point failure. It enables sovereign states to jointly initiate, sustain, and reverse high-risk AI development and deployment at critical thresholds—ensuring controllability, interoperability, and reversibility. By bridging normative principles with enforceable technical mechanisms, this work advances AI safety governance from aspirational guidelines toward operational feasibility, providing a scalable infrastructure for international crisis response and responsible innovation.
Current AI agents in autonomous telecommunications networks lack standardized runtime validation mechanisms, rendering them susceptible to safety risks stemming from erroneous decisions. This work proposes the Guard Rail Validation (GRV) framework, introducing a novel criticality-tiered runtime verification architecture. GRV dynamically assesses decision risk through a weighted evaluation of multidimensional indicators—including action scope, service criticality, autonomy level, reversibility, and temporal behavior—and enforces tiered validation accordingly. The framework supports cross-agent conflict detection, priority arbitration, and compliant audit logging. Evaluated in an O-RAN environment, GRV demonstrates high coverage against known AI/ML attacks and aligns with regulatory requirements such as Article 14 of the EU AI Act.
This study addresses the challenge of evaluating the trustworthiness of automated systems in complex network environments by proposing a five-dimensional “Network Control Intelligence” (NCI) framework. The framework delineates three evolutionary eras of network control and introduces a reference architecture that decouples proposal generation from controlled execution. Emphasizing the synergistic alignment of reasoning capability, verifiability, and authorized execution under large language model (LLM) guidance, it systematically defines, for the first time, the core dimensions of trustworthy autonomous networking. The work not only articulates an integrated paradigm for LLM-enabled network operations and outlines a path toward higher-order autonomy governed by regulatory constraints, but also establishes foundational theoretical principles and design guidelines for secure, governable next-generation network automation.
This study addresses a critical limitation in existing autonomous agent benchmarks, which focus solely on task completion while neglecting the agent’s capacity to proactively abstain from action when faced with missing information, verification, or authorization—thereby inducing compliance bias. To remedy this, the work introduces a novel evaluation paradigm for “abstention capability,” proposing a tripartite taxonomy of abstention scenarios (normative absence, verification absence, and permission absence) and defining three composite metrics: Safety Rate, Usability Rate, and Informed Refusal Rate. Leveraging reinforcement learning from human feedback, the authors analyze the origins of such biases and design a runtime enforcement mechanism for mandatory abstention. Evaluated across 144 enterprise scenarios, the approach achieves up to 89.2% hazardous behavior interception and 87.5% usability in authorized contexts across five model families, demonstrating that the safety–usability trade-off is both tunable and model-dependent.