discrete reward design

Design and analyze finite-valued payoff functions or discrete reward matrices that assign agents rewards from a finite set so as to make specified behavioral outcomes incentive-compatible—e.g., to ensure a target pure strategy profile is a best response and to enforce uniqueness or other equilibrium properties. This competence covers methods for modifying payoff entries from a finite set, constructing general-purpose or task‑agnostic discrete reward schemes, and formulating optimization- or hierarchy-based procedures (optreward) that find discrete rewards satisfying performance, generalization, or constraint requirements.

discreterewarddesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the problem of reconstructing the payoff matrices in two-player games under the constraint that payoffs must belong to a finite discrete set, such that a prescribed pure-strategy profile becomes the unique Nash equilibrium. The work establishes, for the first time, necessary and sufficient feasibility conditions applicable to both zero-sum and general-sum games. Leveraging the discrete nature of the payoff space, the authors devise an efficient dynamic programming algorithm capable of computing exact optimal solutions. In contrast to conventional linear programming–based approaches that assume continuous payoff adjustments, the proposed method guarantees solution exactness while significantly improving computational efficiency, thereby offering a novel paradigm for equilibrium steering in games with discrete payoff constraints.

discrete rewardsequilibrium controlgame changer problem

Recursive Reward Aggregation

Jul 11, 2025
YT
Yuting Tang

In reinforcement learning, designing reward functions for complex objectives often leads to poor behavioral alignment. This paper proposes a novel paradigm that achieves flexible behavioral alignment without modifying the original reward function—instead, it selects an appropriate recursive reward aggregation operator. Grounded in algebraic modeling, we generalize Markov decision processes into a unified framework supporting diverse aggregation operators (e.g., discounted maximum, Sharpe ratio), enabling natural adaptation of the Bellman equation while preserving compatibility with both value-based and actor-critic methods, under both deterministic and stochastic environments. Theoretically, this framework maintains policy iteration convergence guarantees under mild conditions. Empirically, our approach effectively optimizes for heterogeneous complex objectives—including risk-sensitive control, extremum-oriented planning, and long-horizon robustness—demonstrating substantial improvements in alignment generality and practical applicability across benchmark domains.

Aligning agent behavior with complex objectives in RLEliminating reward function modification via aggregationGeneralizing Bellman equations for diverse objectives

This work investigates the reachability of expected reward vectors in multi-objective Markov decision processes (MDPs), particularly addressing the role of randomized policies when pure policies are insufficient. Using tools from convex analysis, probability theory, and MDP theory, the paper establishes that—under any well-defined multidimensional reward structure—every feasible expected reward vector can be approximated to arbitrary precision via a finite convex combination of pure-policy reward vectors; moreover, in the finite-expectation setting, all feasible reward vectors are exactly attainable. The study rigorously characterizes the convex compactness of the payoff set, precisely quantifies the necessity of randomization, and determines the minimal number of pure policies required for such mixtures—reducing policy complexity from infinite to finite mixtures. These results provide foundational theoretical guarantees for designing approximation algorithms in multi-objective MDPs.

Achieving expected payoff vectorsRandomised strategies in MDPsStructure of payoff sets

Automated Design of Affine Maximizer Mechanisms in Dynamic Settings

Feb 12, 2024
MC
Michael Curry
🏛️ Harvard University | University of Zurich | ETH Zurich | Columbia University | Carnegie Mellon University | Optimized Markets | Strategy Robot | Strategic Machine

This paper addresses the challenge of optimizing non-welfare objectives (e.g., revenue) in dynamic mechanism design, where agents’ strategic reporting of private information undermines incentive compatibility. We propose the first general framework that imposes no structural assumptions on valuation functions—such as linearity or monotonicity. Methodologically, we extend affine maximization mechanisms to Markov decision processes (MDPs) with strategic reward reporting, establishing a bi-level optimization paradigm that integrates automated mechanism design and reinforcement learning: the upper level enforces incentive compatibility, while the lower level optimizes the target non-welfare objective. Our contributions are threefold: (1) we eliminate restrictive valuation assumptions, enabling application to arbitrary RL-solvable environments; (2) we automatically synthesize truthful dynamic mechanisms; and (3) our approach significantly outperforms baselines on revenue and other non-welfare objectives, combining theoretical soundness with computational tractability.

Addresses untruthful reward reports in dynamic settingsExtends affine maximizer mechanisms to MDPsOptimizes mechanisms for goals beyond welfare

Latest Papers

What's happening recently
View more

This work addresses the ambiguity and data scarcity inherent in reward function identification for two-player zero-sum games by proposing a unified inverse reward learning framework. Leveraging observed agent policies, the framework reconstructs the underlying reward functions in both entropy-regularized static matrix games and dynamic Markov games. The key innovation lies in establishing, for the first time, identifiability conditions for linear reward functions under quantal response equilibrium, and in designing a general-purpose learning algorithm applicable to both static and dynamic settings. By integrating quantal response equilibrium, entropy regularization, and maximum likelihood estimation, the method achieves sample-efficient learning. Theoretical analysis confirms the algorithm’s reliability and sample efficiency, while numerical experiments demonstrate its effectiveness in competitive decision-making environments.

entropy regularizationinverse reinforcement learningquantal response equilibrium

This work addresses the challenge that non-experts face in systematically constructing alignment reward functions that reflect human preferences. The authors propose a three-step framework: first translating natural language objectives into measurable outcome variables, then modeling the selection of reward terms as a minimum-cost partial set cover problem grounded in a causal graph, and finally iteratively fitting linear reward weights through preference queries. This approach is the first to deterministically characterize the conflict-free feasible weight region and provably converges to a target accuracy within \(O(n \log \kappa)\) queries. By integrating causal reasoning, max-flow computation, and convex feasibility solving, the framework achieves high efficiency, interpretability, and theoretical guarantees, substantially lowering the barrier for non-experts to design aligned reward functions.

human-aligned rewardsobjective decompositionpreference elicitation

This study addresses the challenge of equilibrium nonexistence in multi-principal, multi-team settings, where strategic externalities induce interdependence among incentive-compatible mechanisms and potential discontinuities in the mechanism correspondence. To overcome this limitation of classical models, the authors develop a novel framework that jointly characterizes the outcome distribution along honest obedience paths and the feasible sets attainable through unilateral deviations, integrating mechanism design theory, game theory, and set-valued analysis. Within this framework, they establish rigorous conditions for equilibrium existence in environments featuring team production and agency problems, thereby significantly extending the applicability of Myerson’s classic model to more complex, realistic multi-principal contexts.

equilibrium existenceincentive compatibilitymechanism design

This study addresses the long-standing open problem regarding the computational complexity of optimal contract design over gross substitutes reward function classes in multi-agent binary action models. By conducting a fine-grained analysis of subclasses including OXS, WMRF, and partition rank, combined with approximation algorithm design and randomized query lower bound techniques, this work investigates the inherent tractability boundaries within the gross substitutes class. It reveals that submodularity serves as the critical factor for approximability. Specifically, the authors prove that OXS is APX-complete, establish an EPTAS for WMRF along with an FPTAS under specific settings, and demonstrate the strong inapproximability of ultra rewards. Collectively, these results provide a complete characterization of the complexity landscape for this problem, delineating the precise internal complexity boundaries within gross substitutes for the first time.

approximationcomputational complexitygross substitutes

This work formalizes reward hacking not as a fixable bug but as a structural equilibrium in multi-task principal–agent models, arising when AI systems systematically neglect unmeasured quality dimensions under limited evaluation. Building on five axioms—including multidimensional quality and bounded assessment—the study integrates Holmström–Milgrom agency theory with differentiable reward modeling to derive a computable distortion index that predicts both the direction and severity of reward hacking. The framework further formalizes a “betrayal threshold” mechanism and unifies diverse phenomena such as sycophancy and length gaming under a common theoretical lens. It proves that as the number of deployable tools increases, evaluation coverage asymptotically approaches zero while hacking severity grows without bound, and it provides a pre-deployment vulnerability assessment protocol grounded in this analysis.

AI alignmentevaluation coverageGoodhart's law

Hot Scholars

WL

Weiming Lu

Zhejiang University
Natural Language ProcessingLarge Language ModelsAGI
WZ

Wenqi Zhang

Zhejiang University
Language ModelMultimodal LearningEmbodied Agents
AK

Aviral Kumar

Carnegie Mellon University
AIReinforcement Learning
XH

Xiangnan He

University of Science and Technology of China
RecommendationCausalityBig DataInformation Retrieval
ZL

Ziwei Liu

Associate Professor, Nanyang Technological University
Computer VisionMachine LearningComputer Graphics