A Framework for Designing Reward Functions: From Objectives to Features to Human-Aligned Reward Functions

📅 2026-08-12
📈 Citations: 0
✹ Influential: 0
📄 PDF
đŸ€– AI Summary
This work addresses the challenge that non-experts face in systematically constructing alignment reward functions that reflect human preferences. The authors propose a three-step framework: first translating natural language objectives into measurable outcome variables, then modeling the selection of reward terms as a minimum-cost partial set cover problem grounded in a causal graph, and finally iteratively fitting linear reward weights through preference queries. This approach is the first to deterministically characterize the conflict-free feasible weight region and provably converges to a target accuracy within \(O(n \log \kappa)\) queries. By integrating causal reasoning, max-flow computation, and convex feasibility solving, the framework achieves high efficiency, interpretability, and theoretical guarantees, substantially lowering the barrier for non-experts to design aligned reward functions.
📝 Abstract
We present a formal process to enable non-experts to instantiate and iterate on human-aligned reward functions, i.e. reward functions that adhere to a given preference ordering over trajectories. Given a task described in natural language, our process produces a linear reward function in three steps: distill the task's objectives into a set of fundamental objectives and derive measurable outcome variables that capture those fundamental objectives, select a causally representative subset of outcome variables as the reward terms, and fit weights to those reward terms via preference elicitation. Our contributions describe the first step and formalize the latter two steps. The first is a guided workflow for deriving outcome variables. The second is a reduction of reward term selection to minimum-cost partial cover on a causal DAG, solved in polynomial time via max-flow. The third is a geometric framing of weight fitting as a convex feasibility problem iteratively narrowed by preference queries, solved by existing separation oracle methods. To the best of our knowledge, this is the first reward-design method that maintains a deterministically conflict-free feasible weight region, narrowed to a desired tolerance via a separation oracle with O(n log Îș) preference queries.
Problem

Research questions and friction points this paper is trying to address.

reward function design
human-aligned rewards
preference elicitation
reinforcement learning
objective decomposition
Innovation

Methods, ideas, or system contributions that make the work stand out.

reward function design
causal DAG
preference elicitation
convex feasibility
human-aligned RL
D
Di Yang Shi
University of Texas at Austin
W
W. Bradley Knox
University of Texas at Austin