Score
Designs and documents governance and oversight mechanisms that assign responsibilities and authorities, define reporting, escalation and review processes, and establish monitoring, audit, enforcement, and remediation procedures. Builds and evaluates operational controls, metrics, transparency and feedback loops to measure effectiveness, ensure compliance, and support accountability and continuous improvement.
Current AI governance research lacks systematic integration of diverse frameworks and practices, with notable gaps in the operationalizability of key mechanisms and the implementation of inclusive, stakeholder-centered approaches. To address this, we conduct a rapid three-tier literature review, systematically synthesizing nine authoritative IEEE/ACM reviews published between 2020 and 2024. We introduce the novel “thematic semantic synthesis” analytical paradigm to identify high-frequency governance frameworks (e.g., the EU AI Act, NIST AI Risk Management Framework), core principles (e.g., transparency, accountability), and stakeholder role distributions. Our analysis reveals four critical knowledge gaps in AI governance scholarship and practice. Based on these findings, we propose a rigorously grounded, organizationally feasible governance roadmap—bridging theoretical advancement and real-world implementation. This work contributes both empirical evidence and methodological innovation to advance AI governance research and practice.
This paper identifies a core dilemma in organizational responsible AI governance: ambiguous responsibility boundaries across AI lifecycle stages and a lack of role- and stage-appropriate operational tools. Methodologically, the study systematically reviews over 220 responsible AI tools and proposes a novel two-dimensional (Actor, Stage) classification framework, integrating systematic review, meta-analysis, and qualitative coding. It identifies three critical governance gaps: (1) unclear accountability attribution, (2) absence of empirical validation for most tools, and (3) severe coverage imbalance across actors and stages. Results show that >80% of tools target developers during data and modeling phases; tools for leadership, deployers, end users, and stages such as value proposition definition and deployment are virtually absent. Moreover, >90% of tools lack empirical evidence. The study establishes a theoretically grounded, empirically benchmarked framework to advance actor–stage–aligned AI governance tool ecosystems.
This paper addresses the regulatory failure exacerbated by deploying autonomous AI—particularly embodied agents—in the public sector, where traditional siloed, stage-gated approval mechanisms fail to meet three emerging needs: continuous oversight, deep integration of governance into operational workflows, and cross-agency coordination. Adopting a mixed-methods approach—systematic literature review complemented by in-depth interviews with frontline public officials—the study identifies, for the first time, five core AI governance dimensions tailored to public-sector contexts: cross-agency implementation, holistic assessment, enhanced security, operational transparency, and systemic auditing. Based on these, it proposes a novel “agent-oriented regulatory framework” that is institutionally adaptive and technically interoperable. The framework bridges theory and practice, offering actionable guidance for governing autonomous AI systems under real-world institutional constraints—thereby filling a critical gap in the literature on public-sector AI regulation.
This study addresses the lack of systematic comparative analysis in business process compliance monitoring, particularly for non-conformance checking techniques. Through a systematic literature review (SLR), process mining, compliance modeling, and qualitative comparative analysis, it maps real-world applications across domains, operational workflows, technical foundations, and result representations. The analysis identifies key implementation barriers—especially pervasive human dependence and the absence of standardized evaluation criteria. As the first structured survey framework dedicated to non-conformance checking, the study introduces a standardized, multi-dimensional evaluation framework that clarifies commonalities and distinctions across the technical landscape. It further proposes an extensible theoretical pathway and practical guidelines for automated compliance monitoring. This work provides a methodological foundation and strategic direction for both academic research and industrial deployment. (149 words)
This work addresses the risks of generative interfaces in high-stakes settings—such as hallucinations, semantic distortion, bias, and accessibility barriers—that arise from insufficient human oversight and undermine user understanding and control. To mitigate these issues, the paper proposes a “supervision-by-design” architecture that deeply integrates human judgment into the generative pipeline. This framework employs automated risk detection across dimensions including readability, semantic fidelity, factual consistency, and accessibility compliance, coupled with explicit UI controls and a tiered escalation protocol that triggers mandatory human review upon violations. By combining human-in-the-loop (HITL) interventions with human-on-the-loop (HOTL) continuous monitoring, the approach establishes a scalable, verifiable governance loop. The resulting system significantly enhances transparency, reliability, and inclusivity, thereby strengthening accountability and user agency in high-risk human-AI collaborative decision-making.
This paper addresses the conceptual conflation of “oversight” and “control” in AI safety governance, systematically distinguishing their distinct objectives, operational mechanisms, and temporal scopes. Through a critical cross-disciplinary literature review—and integrating insights from Responsible AI maturity models and risk governance theory—it develops a theoretically rigorous yet policy-actionable analytical framework, introducing the first AI Oversight Maturity Model (AI-OMM). The model identifies critical boundary conditions for oversight failure and establishes a structured, conditional system for assessing the feasibility of meaningful human oversight. Key contributions include: (1) clarifying the normative distinction between oversight and control; (2) diagnosing design gaps and contextual limitations in current oversight mechanisms; and (3) providing regulators, auditors, and developers with a practical tool to evaluate oversight effectiveness, detect capability gaps, and guide technical alignment with governance requirements.
Current AI incident governance frameworks lack consistency in defining, categorizing, monitoring, and reporting incidents, which constrains the depth and accuracy of post-deployment failure analysis. This study addresses this gap through a systematic literature review and comparative analysis across multiple governance frameworks, thereby identifying and synthesizing key inconsistencies that span existing mechanisms. The work reveals systemic deficiencies in data collection practices, classification logics, and analytical rigor, and elucidates critical misalignments among core governance components. By clarifying these structural disconnects, the research establishes a theoretical foundation and proposes a coordinated pathway toward a unified, standardized framework for AI incident governance.
This work addresses the absence of standardized, composable oversight infrastructure in current AI deployments, which leads teams to repeatedly build fragmented auditing and monitoring mechanisms. The authors propose a five-layer, six-dimension framework for AI oversight, with a particular focus on formally defining— for the first time—the “normative layer.” This layer translates human intent into executable, traceable, and upgradable machine-checkable norms through six design principles, including elicitable, adversarially aware, and governable specifications. Integrating formal methods, policy languages (e.g., Cedar, OPA), and constitutional AI concepts, the study introduces CARMA, a norm-driven runtime oversight prototype that demonstrates how a single norm can uniformly drive execution, evaluation, and upgrading. The system validates the feasibility of reusing composable oversight components across teams.
This work addresses the inadequacy of existing large language model (LLM) lifecycle frameworks, which predominantly emphasize operational efficiency while lacking explicit support for security-critical activities—such as data provenance, component signing, and access control—and failing to align governance requirements with specific lifecycle phases. The paper proposes the first security-oriented LLM system lifecycle model, structured not by workflow but by security boundaries, organizing 32 phases into four layered pipelines: data, model, distribution, and application, while integrating LLMOps and governance pillars. It uniquely identifies 13 distinct security-critical phases and exposes a structural imbalance wherein regulatory evidence is concentrated at deployment despite pivotal decisions occurring during development. By mapping key standards—including NIST AI RMF, the EU AI Act, and ISO/IEC 42001—the study establishes a phase-to-governance correspondence mechanism, yielding a comprehensive, lifecycle-spanning security analysis framework that offers structured guidance for compliance and secure design.
This study addresses the pervasive lack of structural integrity in current AI governance documents, which often fail to meet critical requirements such as traceability, dynamic re-verification, and objective evidence. To bridge this gap, the work systematically adapts structural governance principles from aviation software certification standards (DO-178C/DO-330) and proposes a novel integrity framework tailored for static AI governance artifacts. The framework introduces three key concepts—“epoch constraints,” “proof surfaces,” and “structural gaps”—and establishes the seven-principle PromptQ system. Structural analysis of mainstream governance documents reveals that 37% fall below a basic quality threshold, thereby demonstrating the framework’s effectiveness and practicality in enhancing the rigor and verifiability of AI governance documentation.