Score
Designs and conducts assessments that map governance protocols and institutional arrangements to a governance taxonomy, identify and classify capability and interoperability gaps (e.g., supported/partial/absent), and distinguish gaps that are extensible versus structural. Evaluates the time-sensitivity and likely evolution of identified gaps to prioritize remediation and track changes over time.
Existing agent interoperability protocols primarily focus on task coordination and lack support for governance-constrained collective decision-making in multi-agent communities. Drawing on organizational theory and corporate governance standards, this work proposes a six-dimensional governance framework encompassing membership management, deliberation, voting, dissent reservation, human escalation, and audit replay. The study systematically evaluates five prominent protocols—MCP, A2A, ACP, and others—and reveals, for the first time, that the core deficiency lies not in insufficient protocol-level features but in the absence of a dedicated architectural layer for governance. It further distinguishes between scalability gaps and structural gaps. The analysis demonstrates that current protocols universally lack voting and dissent reservation mechanisms, provide only partial support for deliberation, and offer no complete set of governance primitives.
This study addresses the lack of a systematic evaluation framework for AI governance prompts, which undermines their structural integrity as enforceable norms. To bridge this gap, the work proposes an integrative five-principle assessment framework grounded in computability theory, proof theory, and Bayesian epistemology. The authors conduct a static analysis of 34 AGENTS.md files from GitHub, revealing that 37% of file–model pairs fail to meet the defined threshold for structural integrity. Common deficiencies include missing data categorization and absent evaluation criteria, exposing undocumented gaps in artifact classification. These findings provide both theoretical grounding and empirical evidence for developing automated tools capable of detecting and repairing such deficiencies, thereby advancing the formalization and operationalization of AI governance prompts.
Current AI incident governance frameworks lack consistency in defining, categorizing, monitoring, and reporting incidents, which constrains the depth and accuracy of post-deployment failure analysis. This study addresses this gap through a systematic literature review and comparative analysis across multiple governance frameworks, thereby identifying and synthesizing key inconsistencies that span existing mechanisms. The work reveals systemic deficiencies in data collection practices, classification logics, and analytical rigor, and elucidates critical misalignments among core governance components. By clarifying these structural disconnects, the research establishes a theoretical foundation and proposes a coordinated pathway toward a unified, standardized framework for AI incident governance.
In multi-stakeholder platforms, software architecture decisions often implicitly entrench conflicting requirements without systematic support for mapping governance principles to technical design. This work proposes the first governance-architecture alignment framework, explicitly linking five core governance principles to the space of architectural decisions, thereby rendering implicit governance stances identifiable and contestable. The framework also exposes how default technical choices can obscure underlying value commitments. Feasibility is preliminarily demonstrated through a constructive case study of a pig-farming knowledge platform in Rwanda. Future work will employ pre- and post-intervention user judgment studies to evaluate the framework’s impact on actual governance outcomes.
This study addresses the governance challenges organizations face when deploying AI, which stem from an “AI assessability gap”—the lack of sufficient evidence to support high-confidence decisions. The work establishes evidence adequacy as a distinct dimension of AI governance and introduces the concept of “assessability”: the system’s capacity to continuously generate, maintain, and update adequate evidence for governance decisions. It distinguishes between operational and investment certification mechanisms and develops a formal theoretical framework grounded in a confidence function Conf(D|E), integrating structural and causal evidence analysis. Within this framework, six key attributes of assessable evidence are rigorously defined. By providing both theoretical foundations and practical pathways to bridge the assessability gap, this research constitutes a prerequisite for effectively managing AI-related risks and sustainably realizing its value.
This study addresses the pervasive lack of structural integrity in current AI governance documents, which often fail to meet critical requirements such as traceability, dynamic re-verification, and objective evidence. To bridge this gap, the work systematically adapts structural governance principles from aviation software certification standards (DO-178C/DO-330) and proposes a novel integrity framework tailored for static AI governance artifacts. The framework introduces three key concepts—“epoch constraints,” “proof surfaces,” and “structural gaps”—and establishes the seven-principle PromptQ system. Structural analysis of mainstream governance documents reveals that 37% fall below a basic quality threshold, thereby demonstrating the framework’s effectiveness and practicality in enhancing the rigor and verifiability of AI governance documentation.
Current evaluations of AI governance proposals often fall into binary oppositions, overlooking implicit value trade-offs and lacking transparent analytical tools. This work proposes a multidimensional policy analysis framework that integrates expert interviews with computational text analysis to construct an interpretable scoring system across policy attributes, enabling cross-proposal comparison through visualization. Its novelty lies in three aspects: first, a multidimensional evaluation approach that avoids predetermined conclusions and explicitly reveals inherent trade-offs; second, a transparent hybrid methodology combining qualitative expert insights with quantitative computational validation; and third, the introduction of a domain-calibrated model as a benchmark against general-purpose large language models. The framework enables comparable, interpretable assessments of AI governance proposals across multiple attributes, allowing stakeholders to evaluate proposal relevance and coherence according to their own normative priorities.
This work addresses the limitations of existing AI governance frameworks, which rely on static metrics and post-hoc audits and thus lack the capacity for dynamic, real-time assessment of deployment readiness in high-risk systems—particularly regarding fairness discrepancies, threshold sensitivity, and remediation progress. To bridge this gap, the paper proposes the Operational AI Deployment Assurance (OADA) framework, which uniquely models governance uncertainty as an operational challenge within the deployment pipeline. OADA introduces mechanisms such as deployment assurance scores, readiness categorization, threshold stability zones, and governance escalation states to enable closed-loop, dynamic governance from evaluation to deployment. By integrating the Fairness Discrepancy Index (FDI) and FairRisk-FDI with threshold sensitivity analysis and repair-aware assurance evolution, OADA successfully identifies models deemed “compliant” by conventional metrics yet operationally unstable, offering a scalable deployment assurance paradigm for high-stakes domains like medical AI.
This work addresses the absence of standardized, composable oversight infrastructure in current AI deployments, which leads teams to repeatedly build fragmented auditing and monitoring mechanisms. The authors propose a five-layer, six-dimension framework for AI oversight, with a particular focus on formally defining— for the first time—the “normative layer.” This layer translates human intent into executable, traceable, and upgradable machine-checkable norms through six design principles, including elicitable, adversarially aware, and governable specifications. Integrating formal methods, policy languages (e.g., Cedar, OPA), and constitutional AI concepts, the study introduces CARMA, a norm-driven runtime oversight prototype that demonstrates how a single norm can uniformly drive execution, evaluation, and upgrading. The system validates the feasibility of reusing composable oversight components across teams.