🤖 AI Summary
Constitutional AI (CAI) faces two core challenges: ambiguous principle efficacy and difficulty in empirically assessing model adherence. This paper introduces the C3AI framework—the first systematic approach to close the loop between constitutional principle construction and empirical validation. Methodologically, it integrates AI ethics and cognitive psychology principles, employs graph-structured principle curation, and designs a fine-grained constitutional compliance evaluation suite. Key findings reveal that positively framed behavioral principles better align with human preferences, whereas current models exhibit strong compliance with negatively framed prohibitions but significant gaps on positive directives. Post-optimization, constitutional guidance enhances safety without compromising general reasoning capabilities. Crucially, this work establishes—empirically and for the first time—that principle phrasing critically determines human-AI alignment outcomes. It thereby lays both theoretical and practical foundations for verifiable, scalable alignment paradigms.
📝 Abstract
Constitutional AI (CAI) guides LLM behavior using constitutions, but identifying which principles are most effective for model alignment remains an open challenge. We introduce the C3AI framework ( extit{Crafting Constitutions for CAI models}), which serves two key functions: (1) selecting and structuring principles to form effective constitutions before fine-tuning; and (2) evaluating whether fine-tuned CAI models follow these principles in practice. By analyzing principles from AI and psychology, we found that positively framed, behavior-based principles align more closely with human preferences than negatively framed or trait-based principles. In a safety alignment use case, we applied a graph-based principle selection method to refine an existing CAI constitution, improving safety measures while maintaining strong general reasoning capabilities. Interestingly, fine-tuned CAI models performed well on negatively framed principles but struggled with positively framed ones, in contrast to our human alignment results. This highlights a potential gap between principle design and model adherence. Overall, C3AI provides a structured and scalable approach to both crafting and evaluating CAI constitutions.