Score
Designs and evaluates algorithms and decision policies that use multi-armed bandit methods to select among candidate contracts or local obligations for agents or components. These systems adaptively explore and exploit a library of contracts to maximize cumulative or team reward while respecting safety/constraint requirements and enabling decentralized operation without centralized runtime control.
This paper studies optimal contract design in a principal–agent framework where the agent sequentially and adaptively performs multi-step actions: selecting an action, observing its outcome, and dynamically deciding whether to continue or stop, ultimately submitting a realized outcome to trigger payment. It innovatively integrates sequential decision-making and the Pandora’s Box model into contract theory to capture adaptive, information-dependent behavior with multiple attempts. Methodologically, it models the agent’s exploration as an optimal stopping problem under uncertainty, with outcomes drawn from known distributions. Theoretical contributions include: (1) proving that linear contracts are efficiently optimal under outcome independence, yielding a polynomial-time algorithm; (2) providing a polynomial-time optimal algorithm for general contracts when the number of possible outcomes is fixed; and (3) establishing computational hardness by showing that constant-factor approximation is impossible under outcome correlation, thereby characterizing the complexity boundary of the problem.
This paper studies optimal contract design in sequential exploration: an agent sequentially opens $n$ costly boxes—each containing a prize—and selects one, while the principal commits upfront to a nonnegative payment contract to incentivize the agent and maximize expected utility (prize value minus payment). It bridges contract theory and the Pandora’s Box problem. Methodologically, it integrates game-theoretic modeling, dynamic programming, optimal stopping theory, and probabilistic analysis. The contributions are threefold: (i) the first polynomial-time algorithm for computing the exact optimal linear contract; (ii) a closed-form characterization of the optimal general (nonlinear) contract under single-prize and i.i.d. box assumptions; and (iii) theoretical guarantees of both principal-optimal utility and incentive compatibility. The framework yields a computationally tractable and interpretable mechanism design paradigm for sequential decision-making under uncertainty.
This paper studies a multi-armed bandit (MAB) problem where strategic agents serve as “arms”: each agent can manipulate reported rewards and incurred costs, necessitating an incentive-compatible mechanism that elicits high-performance truthful behavior while ensuring robust performance under non-equilibrium behavior (e.g., irrationality or deviations). To this end, we propose the first MAB framework jointly guaranteeing incentive compatibility and non-equilibrium robustness. We identify a key structural property enabling synergistic achievement of both objectives and integrate insights from second-price auctions to handle settings with no prior knowledge of arm qualities. Theoretically, our algorithm yields a non-vacuous lower bound on cumulative reward under arbitrary agent behavior; moreover, even without knowledge of true arm performances, it achieves an $O(sqrt{T})$ regret upper bound—substantially improving upon conventional approaches that either ignore incentives or lack robustness guarantees.
This paper studies the online contract design problem, where a principal seeks to maximize utility by learning optimal contracts through multi-round interactions with agents of unknown types (i.e., unknown cost and output functions). We establish the first theoretical equivalence between online contract design and dynamic pricing, Lipschitz multi-armed bandits, and polynomial sample-complexity learning. We show that one-dimensional effort spaces are particularly amenable to learning-driven contract design. Our method leverages nonparametric estimation and regularity analysis to develop efficient algorithms: for binary outcomes, it achieves optimal learning of linear contracts; for homogeneous and heterogeneous agents, it attains near-optimal contracts with provably efficient convergence. We rigorously characterize sample complexity and regret bounds, providing the first systematic theoretical framework for learning-based principal–agent mechanisms.
This paper studies incentive mechanism design in principal-agent multi-armed bandit games, where the agent must autonomously explore an unknown environment—rather than being fully informed and greedy. Addressing the realistic setting where agents are self-interested and exploration is inherently uncertain, we propose the first robust elimination framework tailored to exploration-aware learning agents. Our method integrates adaptive search, robust incentive design, and a coupled Bayesian/frequentist estimation scheme with explicit exploration. Under i.i.d. rewards, it achieves a regret bound of $widetilde{O}(T^{2/3})$; under linear rewards, it attains the optimal $widetilde{O}(sqrt{T})$ bound—improving significantly over Dogan et al. (2023a)’s $widetilde{O}(T^{11/12})$. This provides the first theoretically optimal solution for principals to dynamically guide learning agents in exploration-exploitation settings.
本文通过引入随机性改进了动态委托-代理问题中的激励设计,使用Heteroscedastic GP-UCB算法解决了因不确定性导致的计算难题。
This work addresses the inherent tension in multi-armed bandits between maximizing cumulative reward and accurately estimating the mean rewards of individual arms. The study is the first to systematically characterize the fundamental trade-off between “exploration for learning” and “exploration for profit,” proposing a unified algorithmic framework that flexibly interpolates between these two objectives via an adjustable preference parameter. Theoretical analysis establishes matching upper and lower bounds, demonstrating that the proposed algorithm achieves an optimal balance between regret and estimation error. Empirical evaluations confirm that the method effectively reconciles reward accumulation with estimation accuracy across diverse preference settings.
This work proposes a generalized multi-armed bandit and stopping problem framework grounded in behavioral preferences, departing from conventional modeling paradigms that rely on predefined states, rewards, and transition dynamics. Starting from the decision maker’s preferences over local temporal plans, the authors develop a generalized stopping representation through behavioral axioms and introduce a calendar-time cross-plan pricing mechanism. Under compact time constraints, this approach yields a rested-bandit model exhibiting index optimality. The key innovation lies in interpreting the index as the shadow price of advancing a local clock, thereby unifying—within a single preference-based framework—a diverse array of decision models, including expected utility, learning, robust, rank-dependent, Choquet, and Pandora’s box formulations. This constitutes the first theoretical foundation for index policies rooted entirely in preference theory.
This study investigates the ability of evolutionary algorithms to identify the Condorcet winner within the Dueling Bandits framework. Addressing the exploration–exploitation trade-off inherent in multi-armed bandit settings, the authors employ steady-state analysis of Markov chains to demonstrate, for the first time, that the (1+1) Evolutionary Algorithm (EA) selects the Condorcet winner with only constant probability when the pairwise winning probability satisfies \( p = \Omega(1/n) \). In contrast, an Estimation-of-Distribution Algorithm (EDA) based on MMAS achieves a significantly higher success probability of \( 1 - \Theta(p) \). The work further introduces a repeated dueling mechanism that substantially enhances the Condorcet winner identification performance of EAs.
This work addresses the challenge of ensuring global safety in multi-agent reinforcement learning, where individual agents cannot independently guarantee safety due to action validity depending on the dynamic behaviors of others. The authors propose a contract-based compositional shielding mechanism that provides deterministic safety guarantees under decentralized execution: agents collaboratively select local obligation tuples under a shared global LTL safety specification, balancing team optimality and runtime safety. The approach generates local action masks through verifiable local obligation contracts and dynamically optimizes obligation combinations during end-to-end training using a non-stationary multi-armed bandit algorithm, thereby recovering safety-critical collaborative behaviors that would otherwise require explicit coordination. Evaluated across six environments and fifteen algorithmic variants, this method demonstrates, for the first time, team-optimal safe policies without centralized control.