defeasible policy representation

Design and implement symbolic policy representations that encode normative statements as defeasible rules whose conclusions can be overridden when conflicts arise; explicitly represent priority relations among rules so automated reasoning determines which rules prevail. Build tooling and analyses that produce rule-level explanations of policy-driven decisions and support targeted interventions such as modifying rules or adjusting priorities to alter policy outcomes.

defeasiblepolicyrepresentation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.53
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of ensuring that generative AI systems adhere strictly to written policies when responding to natural language queries, a task where existing approaches struggle to simultaneously achieve accuracy, interpretability, and robustness. The authors propose a hybrid symbolic method that formalizes policies as logical rules, leverages large language models to extract factual predicates from user inputs, and employs an Answer Set Programming (ASP) solver for structured, rule-based reasoning. By decoupling fact extraction from logical inference, this approach significantly enhances robustness against input perturbations while guaranteeing auditability and interpretability of decisions. Experimental results demonstrate that the method outperforms baseline strategies—such as policy-as-prompt and policy-as-code—in accuracy across most scenarios and reduces token consumption by approximately an order of magnitude.

generative AInatural language queriespolicy compliance

Current evaluations of AI governance proposals often fall into binary oppositions, overlooking implicit value trade-offs and lacking transparent analytical tools. This work proposes a multidimensional policy analysis framework that integrates expert interviews with computational text analysis to construct an interpretable scoring system across policy attributes, enabling cross-proposal comparison through visualization. Its novelty lies in three aspects: first, a multidimensional evaluation approach that avoids predetermined conclusions and explicitly reveals inherent trade-offs; second, a transparent hybrid methodology combining qualitative expert insights with quantitative computational validation; and third, the introduction of a domain-calibrated model as a benchmark against general-purpose large language models. The framework enables comparable, interpretable assessments of AI governance proposals across multiple attributes, allowing stakeholders to evaluate proposal relevance and coherence according to their own normative priorities.

AI governancenormative assumptionspolicy analysis

Normative active inference: A numerical proof of principle for a computational and economic legal analytic approach to AI governance

Nov 24, 2025
AC
Axel Constant
🏛️ University of Sussex | VERSES | Université du Québec à Montréal | Sheffield Hallam University

Current AI governance relies heavily on external human oversight, lacking mechanisms for autonomous legal compliance. Method: This paper proposes an “architecture-as-regulation” framework that embeds legal norms directly into AI decision-making by modeling law as a structural scaffold for rational choice under uncertainty. It integrates active inference (AIF), Bayesian updating, Markov decision processes (MDPs), and economic legal analysis (ELA) to construct AI agents with intentional reasoning and normative sensitivity. A novel safety-valve mechanism enables context-dependent preference modeling and dynamic trade-offs among competing objectives. Contribution/Results: Empirical evaluation in autonomous vehicle right-of-way scenarios demonstrates that the framework enables real-time reconciliation of legal constraints with operational goals, significantly improving legal alignment and risk controllability of AI behavior without sacrificing task performance.

Developing regulation by design for autonomous agent decision-makingModeling legal norm influence on AI behavior through active inferenceResolving legal-pragmatic conflicts via context-dependent preference mechanisms

Automated welfare eligibility systems frequently produce explanations misaligned with statutory requirements, undermining administrative legitimacy. This paper introduces an accountable neuro-symbolic framework for public-sector AI—first achieving semantic alignment and formal verifiability between algorithmic explanations and California’s CalFresh statutory law. Methodologically, the framework integrates structured legal ontology modeling, a rule extraction pipeline, formal representation in Prolog and Answer Set Programming (ASP), and solver-driven compliance reasoning. It precisely detects unlawful explanations, pinpoints violated statutory provisions, and enables explanation provenance tracking and contestation—thereby substantially enhancing decision transparency and statutory consistency. The core contribution is a legally traceable explanatory infrastructure that bridges AI interpretability with statutory compliance in public welfare administration.

Develops a legally grounded explainability framework for AI decisionsEvaluates explanations for legal consistency and supports procedural accountabilityLinks automated eligibility justifications to statutory constraints in public benefits

This work proposes GRACE, a novel reason-based neurosymbolic architecture designed to address the challenge of aligning highly autonomous AI systems with both operational efficiency and diverse ethical norms in real-world applications. GRACE decouples moral reasoning from instrumental decision-making through three integrated modules—moral, decision, and guardian—enabling coexistence of multiple ethical frameworks, traceable behavior, and formal verifiability. By combining deontic logic, neurosymbolic reasoning, and large language models, GRACE implements an interpretable, contestable, and formally verifiable alignment mechanism, demonstrated in a psychotherapy assistant case study. This approach empowers stakeholders to understand, challenge, and refine AI behavior while providing dual guarantees of statistical and formal alignment.

AI alignmentautonomous agentsethical AI

Latest Papers

What's happening recently
View more

This study addresses the challenge of AI systems misinterpreting or excessively refusing requests due to implicitly conflicting linguistic rules. To resolve this interpretability bottleneck, it introduces the concept of "consequence graphs," which explicitly link agents, authorities, and policy conditions. Methodologically, compliance judgments are modeled via semantic geometric vectors, with systematic validation conducted through lexical routing techniques and synthetic experiments. The findings reveal that geometric approaches fail to achieve reliable pre-action compliance assessment, and that single retrieval strategies are insufficient for complex decision-making tasks. Ultimately, this work exposes fundamental limitations in existing representation paradigms under rule-conflict scenarios, providing critical corrective directions for future alignment research.

AI judgementautonomous agentspolicy conflict resolution

This work addresses critical challenges in safety-critical rule-based systems—namely poor scalability, fragility, and goal mis-specification—which often lead to reward hacking and failures in formal verification. To overcome these limitations, the authors propose a neuro-symbolic causal framework that integrates first-order logic abductive trees, structural causal models, and deep reinforcement learning within a MAPE-K control loop. A novel meta-layer architecture enables the automatic synthesis and formal verification of rules from natural language objectives. This meta-layer comprises a goal/rule synthesizer and a rule verification engine, which iteratively generate necessary and sufficient causal rule sets grounded in legal and safety principles provided by human experts. Evaluated in an autonomous driving scenario, the approach successfully derives a minimal yet complete rule set, formally encoded as logical constraints, demonstrating its modularity, traceability, and practical applicability.

formal verificationgoal misspecificationreward hacking

Existing AI planning approaches struggle to effectively model dynamic and potentially conflicting social norms, limiting their ability to generate norm-compliant behaviors in human–AI interaction. This work proposes a novel method that integrates defeasible reasoning with norm-guided planning, embedding dynamic social norms as retractable constraints within the planning process to enable real-time, norm-aware decision-making. The approach synergistically combines defeasible logical inference, norm-driven planning, and formal verification, and is implemented within a natural language dialogue system. Theoretical analysis establishes its correctness, while experiments with SocialBot demonstrate that the AI can generate responses that are both contextually appropriate and compliant with evolving social norms—marking the first realization of real-time adherence to and resolution of conflicts among dynamic social norms in human–AI interaction.

AI planningdynamically changing normshuman-AI interaction

Current regulatory decisions often undermine fairness and legitimacy due to their static nature, lack of interpretability, and susceptibility to dominant interest groups. This work proposes a regulatory recommendation system that integrates distributed artificial intelligence with value-sensitive design, enabling independent modeling of diverse stakeholder preferences and their dynamic aggregation through interpretable and verifiable mechanisms. By adaptively responding to evolving factual and normative contexts, the approach uniquely combines distributed AI with value-sensitive principles to significantly enhance the transparency, equity, and social acceptability of regulatory recommendations. Consequently, it strengthens the perceived legitimacy, compliance, and public trust in regulatory processes.

explainabilityjusticelegitimacy

This study addresses the challenge that existing computational methods struggle to capture the implicit argumentative structures embedded in policy texts, where deliberative and managerial discourses are intricately intertwined. To overcome this limitation, the authors propose a bipolar argumentation framework that integrates large language models with symbolic reasoning. The approach first employs a large language model to classify argument frames and then applies deterministic rules to identify four types of mediating relationships: agency attenuation, agenda shifting, instrumental support, and normative support. These relationships are formalized for the first time as computable subtypes, enabling the construction of policy argumentation graphs that are both expert-verifiable and stable across jurisdictions. The work introduces the first annotated dataset of 100 subdocuments drawn from disaster risk reduction policies in the United States, United Kingdom, Canada, and Australia, demonstrating that the resulting graphs achieve high accuracy, interpretability, and cross-domain consistency.

argumentationcritical discourse analysisdisaster governance

Hot Scholars

JP

Jay Pujara

Research Associate Professor, University of Southern California
Machine LearningKnowledge GraphsEfficient Prediction
GH

Georges Hattab

Adjunct Professor of Computer Science, Freie Universität Berlin, Robert Koch Institute
Artificial IntelligenceData MiningVisualization
AD

Akshat Dubey

Research Associate at Robert Koch Institute, Berlin and PhD at Freie Universität, Berlin, Germany
AI RegulationsExplainable AIHuman-Centered AIDeep Learning
JM

Jonathan May

University of Southern California, Information Sciences Institute
Machine TranslationMachine LearningNatural Language Processing
MK

Mayank Kejriwal

University of Southern California
Knowledge GraphsComplex SystemsAI for Social GoodComputational Social Science