From Constitutions to Control: Interpretable Rewards for Aligning Language Models

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the opacity and limited steerability of reward signals in existing alignment methods by proposing a rubric-based framework that pioneers a mapping mechanism from general constitutions to controllable rewards, translating constitutional principles into interpretable and tunable reward models. Methodologically, it employs constitution-guided AI feedback to estimate initial weights, combined with rubric decomposition and multi-dimensional reward reweighting algorithms during training to enable independent modulation of specific behaviors. Experimental results demonstrate that this approach predictively alters target behaviors, effectively mitigates sycophancy and statistical biases, resolves safety-utility trade-offs, and significantly enhances the controllability and transparency of language model alignment.
📝 Abstract
Current approaches to aligning language models often make it hard to know what behavior is being rewarded or to change that reward in a targeted way. In particular, standard preference-based methods collapse multiple considerations into aggregate human judgments, obscuring what drives the resulting reward, while principle-based methods specify high-level values without fully operationalizing them. To address this gap, we develop a rubric-based framework to transform a general-purpose constitution into an interpretable and tunable reward model, using constitution-guided AI feedback to estimate initial weights for the constituent rubric items. We then reweight those dimensions to construct modified rewards for training. Across experiments on political alignment and safety-helpfulness tradeoffs, reweighting individual dimensions predictably changes targeted behaviors largely independently while navigating tradeoffs between conflicting alignment objectives. We show that the same framework can mitigate label bias encoded in preference judgments -- including sycophancy and demographic bias -- by reducing their influence on the training reward. Our results demonstrate that constitution-derived, interpretable rewards can translate high-level alignment principles into more transparent and controllable model behavior.
Problem

Research questions and friction points this paper is trying to address.

language model alignment
interpretable rewards
preference-based methods
label bias
controllable behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

interpretable rewards
rubric-based framework
constitution-guided alignment
reward reweighting
label bias mitigation
🔎 Similar Papers