Score
Design and implement symbolic policy representations that encode normative statements as defeasible rules whose conclusions can be overridden when conflicts arise; explicitly represent priority relations among rules so automated reasoning determines which rules prevail. Build tooling and analyses that produce rule-level explanations of policy-driven decisions and support targeted interventions such as modifying rules or adjusting priorities to alter policy outcomes.
This work addresses the challenge of ensuring that generative AI systems adhere strictly to written policies when responding to natural language queries, a task where existing approaches struggle to simultaneously achieve accuracy, interpretability, and robustness. The authors propose a hybrid symbolic method that formalizes policies as logical rules, leverages large language models to extract factual predicates from user inputs, and employs an Answer Set Programming (ASP) solver for structured, rule-based reasoning. By decoupling fact extraction from logical inference, this approach significantly enhances robustness against input perturbations while guaranteeing auditability and interpretability of decisions. Experimental results demonstrate that the method outperforms baseline strategies—such as policy-as-prompt and policy-as-code—in accuracy across most scenarios and reduces token consumption by approximately an order of magnitude.
Current evaluations of AI governance proposals often fall into binary oppositions, overlooking implicit value trade-offs and lacking transparent analytical tools. This work proposes a multidimensional policy analysis framework that integrates expert interviews with computational text analysis to construct an interpretable scoring system across policy attributes, enabling cross-proposal comparison through visualization. Its novelty lies in three aspects: first, a multidimensional evaluation approach that avoids predetermined conclusions and explicitly reveals inherent trade-offs; second, a transparent hybrid methodology combining qualitative expert insights with quantitative computational validation; and third, the introduction of a domain-calibrated model as a benchmark against general-purpose large language models. The framework enables comparable, interpretable assessments of AI governance proposals across multiple attributes, allowing stakeholders to evaluate proposal relevance and coherence according to their own normative priorities.
Current AI governance relies heavily on external human oversight, lacking mechanisms for autonomous legal compliance. Method: This paper proposes an “architecture-as-regulation” framework that embeds legal norms directly into AI decision-making by modeling law as a structural scaffold for rational choice under uncertainty. It integrates active inference (AIF), Bayesian updating, Markov decision processes (MDPs), and economic legal analysis (ELA) to construct AI agents with intentional reasoning and normative sensitivity. A novel safety-valve mechanism enables context-dependent preference modeling and dynamic trade-offs among competing objectives. Contribution/Results: Empirical evaluation in autonomous vehicle right-of-way scenarios demonstrates that the framework enables real-time reconciliation of legal constraints with operational goals, significantly improving legal alignment and risk controllability of AI behavior without sacrificing task performance.
Automated welfare eligibility systems frequently produce explanations misaligned with statutory requirements, undermining administrative legitimacy. This paper introduces an accountable neuro-symbolic framework for public-sector AI—first achieving semantic alignment and formal verifiability between algorithmic explanations and California’s CalFresh statutory law. Methodologically, the framework integrates structured legal ontology modeling, a rule extraction pipeline, formal representation in Prolog and Answer Set Programming (ASP), and solver-driven compliance reasoning. It precisely detects unlawful explanations, pinpoints violated statutory provisions, and enables explanation provenance tracking and contestation—thereby substantially enhancing decision transparency and statutory consistency. The core contribution is a legally traceable explanatory infrastructure that bridges AI interpretability with statutory compliance in public welfare administration.
This work proposes GRACE, a novel reason-based neurosymbolic architecture designed to address the challenge of aligning highly autonomous AI systems with both operational efficiency and diverse ethical norms in real-world applications. GRACE decouples moral reasoning from instrumental decision-making through three integrated modules—moral, decision, and guardian—enabling coexistence of multiple ethical frameworks, traceable behavior, and formal verifiability. By combining deontic logic, neurosymbolic reasoning, and large language models, GRACE implements an interpretable, contestable, and formally verifiable alignment mechanism, demonstrated in a psychotherapy assistant case study. This approach empowers stakeholders to understand, challenge, and refine AI behavior while providing dual guarantees of statistical and formal alignment.
This study addresses the challenge of AI systems misinterpreting or excessively refusing requests due to implicitly conflicting linguistic rules. To resolve this interpretability bottleneck, it introduces the concept of "consequence graphs," which explicitly link agents, authorities, and policy conditions. Methodologically, compliance judgments are modeled via semantic geometric vectors, with systematic validation conducted through lexical routing techniques and synthetic experiments. The findings reveal that geometric approaches fail to achieve reliable pre-action compliance assessment, and that single retrieval strategies are insufficient for complex decision-making tasks. Ultimately, this work exposes fundamental limitations in existing representation paradigms under rule-conflict scenarios, providing critical corrective directions for future alignment research.
This work addresses critical challenges in safety-critical rule-based systems—namely poor scalability, fragility, and goal mis-specification—which often lead to reward hacking and failures in formal verification. To overcome these limitations, the authors propose a neuro-symbolic causal framework that integrates first-order logic abductive trees, structural causal models, and deep reinforcement learning within a MAPE-K control loop. A novel meta-layer architecture enables the automatic synthesis and formal verification of rules from natural language objectives. This meta-layer comprises a goal/rule synthesizer and a rule verification engine, which iteratively generate necessary and sufficient causal rule sets grounded in legal and safety principles provided by human experts. Evaluated in an autonomous driving scenario, the approach successfully derives a minimal yet complete rule set, formally encoded as logical constraints, demonstrating its modularity, traceability, and practical applicability.
Existing AI planning approaches struggle to effectively model dynamic and potentially conflicting social norms, limiting their ability to generate norm-compliant behaviors in human–AI interaction. This work proposes a novel method that integrates defeasible reasoning with norm-guided planning, embedding dynamic social norms as retractable constraints within the planning process to enable real-time, norm-aware decision-making. The approach synergistically combines defeasible logical inference, norm-driven planning, and formal verification, and is implemented within a natural language dialogue system. Theoretical analysis establishes its correctness, while experiments with SocialBot demonstrate that the AI can generate responses that are both contextually appropriate and compliant with evolving social norms—marking the first realization of real-time adherence to and resolution of conflicts among dynamic social norms in human–AI interaction.
Current regulatory decisions often undermine fairness and legitimacy due to their static nature, lack of interpretability, and susceptibility to dominant interest groups. This work proposes a regulatory recommendation system that integrates distributed artificial intelligence with value-sensitive design, enabling independent modeling of diverse stakeholder preferences and their dynamic aggregation through interpretable and verifiable mechanisms. By adaptively responding to evolving factual and normative contexts, the approach uniquely combines distributed AI with value-sensitive principles to significantly enhance the transparency, equity, and social acceptability of regulatory recommendations. Consequently, it strengthens the perceived legitimacy, compliance, and public trust in regulatory processes.
This study addresses the challenge that existing computational methods struggle to capture the implicit argumentative structures embedded in policy texts, where deliberative and managerial discourses are intricately intertwined. To overcome this limitation, the authors propose a bipolar argumentation framework that integrates large language models with symbolic reasoning. The approach first employs a large language model to classify argument frames and then applies deterministic rules to identify four types of mediating relationships: agency attenuation, agenda shifting, instrumental support, and normative support. These relationships are formalized for the first time as computable subtypes, enabling the construction of policy argumentation graphs that are both expert-verifiable and stable across jurisdictions. The work introduces the first annotated dataset of 100 subdocuments drawn from disaster risk reduction policies in the United States, United Kingdom, Canada, and Australia, demonstrating that the resulting graphs achieve high accuracy, interpretability, and cross-domain consistency.