Score
Designs and evaluates methods to extract or synthesize concise, human‑interpretable "constitutions"—sets of steering principles or decision rules—that summarize model preferences, pairwise choices, or rationale texts. This includes techniques for inverse constitutional AI (reconstructing rules from model behavior), compressing preference datasets into principles, aggregating debate rationales into constitutions, and democratic or ICAI approaches to produce compact, interpretable guidance.
Existing constitutional preference elicitation methods represent principles as flat lists, lacking compositional logic, which renders decision rules unexecutable and inadequately evaluated. This work systematically investigates the performance of constitutional approaches on three open challenges: principle quality metrics, compositional ambiguity, and inter-model variability. It empirically demonstrates for the first time that constitutions must be evaluated as integrated “constitution–executor” systems. To address these limitations, the paper introduces ICAI+, a principled refinement framework. Experiments on PRISM, AlpacaEval, and Chatbot Arena datasets—employing executors such as Inverse Constitutional AI (ICAI), LLM judges, and majority voting—show that ICAI+ improves agreement across executors from 73% to 78%, while achieving comparable accuracy between transparent executors (66%) and LLM judges (67%).
This work addresses the problem of interpretable modeling for pairwise text preference data. We propose the Inverse Constitutional AI (ICAI) paradigm, which formalizes preference explanation as a “constitution compression” task—automatically distilling concise, verifiable preference principles from human-annotated feedback. Methodologically, ICAI integrates large language model–driven iterative prompt optimization with a constitution distillation algorithm, validated via synthetic data evaluation, cross-annotator human assessment, and multi-group PRISM empirical analysis. Our key contribution is the first formulation of opaque preference feedback as a transparent, learnable, and generalizable constitution generation problem. Experiments demonstrate near-perfect (≈100%) preference reconstruction accuracy on synthetic data and statistically significant improvements over baselines on real-world benchmarks—including AlpacaEval, Chatbot Arena, and PRISM. The generated constitutions exhibit both bias-detection capability and strong cross-sample generalizability.
This work addresses the limitations of existing preference alignment methods, which struggle to capture the multidimensional reasoning underlying human judgments and rely solely on pairwise labels, thereby constraining interpretability and expressiveness. To overcome these challenges, the authors propose a structured role-based debate framework for preference modeling: multiple perspectives engage in competitive debate to generate rationales, which are then distilled into natural language guiding principles. These principles are integrated with large language model prompting and decision trees to predict preferences. The approach substantially enhances both the expressiveness of the derived principles and the interpretability of decisions. Evaluated on the MuCE-Pref and LiTBench benchmarks, the method outperforms current baselines in preference prediction accuracy, and its generated principles are consistently preferred by human evaluators.
Large language models (LLMs) suffer from limited interpretability and implicit alignment principles, leading to transparency and consistency bottlenecks in alignment. Method: This paper proposes Enhanced Inverse Constitutional AI—a novel framework that systematically optimizes principle generation, semantic clustering, and contrastive embedding learning to explicitly extract high-quality, generalizable alignment principles from preference data, replacing opaque, implicit alignment. It introduces a principle refinement mechanism enabling joint modeling of synthetic and real-world data, supporting cross-domain generalization and auditable verification. Contribution/Results: The extracted principles exhibit strong interpretability and stability. While contextual alignment gains are modest, the method establishes a critical foundation for a new alignment paradigm: tuning-free, plug-and-play, and verifiable—enabling transparent, principled, and auditable LLM alignment.
Constitutional AI (CAI) faces two core challenges: ambiguous principle efficacy and difficulty in empirically assessing model adherence. This paper introduces the C3AI framework—the first systematic approach to close the loop between constitutional principle construction and empirical validation. Methodologically, it integrates AI ethics and cognitive psychology principles, employs graph-structured principle curation, and designs a fine-grained constitutional compliance evaluation suite. Key findings reveal that positively framed behavioral principles better align with human preferences, whereas current models exhibit strong compliance with negatively framed prohibitions but significant gaps on positive directives. Post-optimization, constitutional guidance enhances safety without compromising general reasoning capabilities. Crucially, this work establishes—empirically and for the first time—that principle phrasing critically determines human-AI alignment outcomes. It thereby lays both theoretical and practical foundations for verifiable, scalable alignment paradigms.
This work addresses the limitations of existing constitutional learning methods for large language models, which rely heavily on extensive annotated data and unstructured prompts that hinder scalability. The authors propose MAC (Multi-Agent Constitutional learning) and its enhanced variant MAC+, a framework employing collaborative agents—each responsible for accepting, editing, or rejecting rule updates—to iteratively refine a human-readable and auditable set of natural language rules. By integrating reinforcement learning with trajectory replay, the approach efficiently learns behavioral policies without requiring model parameter updates. Evaluated on low-resource tasks such as PII annotation, MAC and MAC+ outperform current prompt optimization techniques by over 50% and achieve performance comparable to supervised fine-tuning and GRPO, demonstrating both efficacy and scalability in resource-constrained settings.
This work addresses the absence of end-to-end polynomial-time procedures in existing autonomous governance mechanisms, where the core aggregation problem is NP-hard. The authors propose a constitutional governance framework that unifies voting, proposal formation, deliberation, amendment, and consensus into an efficient autonomous process: members submit ideal proposals, which are synthesized—via coalition formation, AI-mediated negotiation, and supermajority support—into public proposals scored and adopted according to constitutional rules. The framework innovatively integrates metric space aggregation centered on the generalized median, reality-aware social choice, supermajority-based constitutional amendment, and self-amending constitutional mechanisms. Theoretically, sincere voting is shown to weakly dominate strategic misreporting; the compromise gap vanishes in one-dimensional settings and remains bounded in general cases. Empirical validation across seven instance types demonstrates the framework’s efficacy, with simulations indicating that proposed heuristics substantially reduce the compromise gap.
Current evaluations of AI governance proposals often fall into binary oppositions, overlooking implicit value trade-offs and lacking transparent analytical tools. This work proposes a multidimensional policy analysis framework that integrates expert interviews with computational text analysis to construct an interpretable scoring system across policy attributes, enabling cross-proposal comparison through visualization. Its novelty lies in three aspects: first, a multidimensional evaluation approach that avoids predetermined conclusions and explicitly reveals inherent trade-offs; second, a transparent hybrid methodology combining qualitative expert insights with quantitative computational validation; and third, the introduction of a domain-calibrated model as a benchmark against general-purpose large language models. The framework enables comparable, interpretable assessments of AI governance proposals across multiple attributes, allowing stakeholders to evaluate proposal relevance and coherence according to their own normative priorities.
Existing approaches struggle to discern which constitutional values language models genuinely prioritize in value conflicts, as reliance solely on output behavior often leads to misjudgment. This work proposes the Constitutional Value Potential (CVP) framework, which, for the first time, learns scalar potentials representing individual constitutional values directly from model hidden states. By capturing priority-margin signals in activation space, CVP enables interpretable monitoring and targeted intervention of a model’s intrinsic value orientations. The method integrates independent adjudicator supervision, hidden-state probing, and directional intervention tests, demonstrating strong efficacy on the Qwen2.5 model series: the monitor achieves 0.95 AUROC in predicting value-conflict violations, significantly outperforming strong baselines and generalizing to unseen synthetic conflicts. Moreover, early activation signals reliably indicate whether adversarial priority attacks have succeeded.