π€ AI Summary
This study addresses the limitations of traditional models in explaining the evolutionary mechanisms of fairness perception, which often overlook the multidimensionality of human decision-making. To this end, we propose a multi-objective reinforcement learning framework that introduces a fairness pressure coefficient and employs a dual-objective Q-learning algorithm to simulate the Ultimatum Game, capturing the dynamic trade-off between material payoffs and moral fairness. Our simulations reveal a reversal phenomenon of "tolerant" strategies under moderate pressure alongside preference competition mechanisms, validating the critical influence of fairness pressure on decision outcomes. By extending the reinforcement learning paradigm, this work offers a novel analytical perspective for understanding broader social behaviors.
π Abstract
Fairness, as a fundamental social norm, continues to pose a longstanding puzzle regarding its emergence. Traditional game-theoretic models largely rely on the assumption of \emph{Homo economicus}, wherein individuals are purely rational and self-interested, acting solely to maximize material payoffs. Such accounts, however, overlook the multidimensional nature of human decision-making, which is often shaped also by other considerations beyond economic incentives. To address this gap, we propose a multi-objective reinforcement learning framework that models the evolution of fairness as a dynamic trade-off between material payoff maximization and fairness-driven moral behavior, regulated by a fairness pressure coefficient. Using simulations of a two-objective Q-learning ultimatum game, we find that increased fairness pressure promotes fair outcomes, as expected. Strikingly, however, under moderate pressure, responder behavior reverses: responders become ``forgiving"by accepting low offers -- a pattern in line with our daily experience. Microscopic analyses reveal that this strategy reversal stems from competition between payoff-maximizing and fairness-oriented preferences. We further extend our framework to an asymmetric setting, where proposers and responders assign different weights to the two objectives. Overall, our work expands the reinforcement learning paradigm from a single-objective to a multi-objective formulation, offering a versatile tool for elucidating a broader range of human social behaviors.