Score
Designs, builds, or analyzes incentive-aware exploration mechanisms that combine UCB-style bandit selection with a payment-freezing rule: protocols that temporarily fix or suspend payments to agents to stabilize incentives while performing upper-confidence-bound exploration. Work includes specifying the UCB allocation and payment-freezing procedure, proving bounds on welfare regret and incentive error, and tuning freeze/learning parameters to target rates such as ~o(√(ngt)) or ~o(t^{2/3}).
This paper studies incentive mechanism design in principal-agent multi-armed bandit games, where the agent must autonomously explore an unknown environment—rather than being fully informed and greedy. Addressing the realistic setting where agents are self-interested and exploration is inherently uncertain, we propose the first robust elimination framework tailored to exploration-aware learning agents. Our method integrates adaptive search, robust incentive design, and a coupled Bayesian/frequentist estimation scheme with explicit exploration. Under i.i.d. rewards, it achieves a regret bound of $widetilde{O}(T^{2/3})$; under linear rewards, it attains the optimal $widetilde{O}(sqrt{T})$ bound—improving significantly over Dogan et al. (2023a)’s $widetilde{O}(T^{11/12})$. This provides the first theoretically optimal solution for principals to dynamically guide learning agents in exploration-exploitation settings.
Existing incentive mechanisms fail to adapt to learning agents whose strategies evolve continuously over time. Method: We propose a two-timescale adaptive incentive mechanism driven by individual externality—the difference between an agent’s marginal cost and the operator’s marginal cost—and updated at a rate slower than the agents’ learning dynamics. Leveraging two-timescale stochastic approximation and differential game theory, we develop a technical framework comprising externality modeling, fixed-point analysis, and convergence proof. Contribution/Results: We establish the first general incentive framework decoupled from agents’ learning dynamics; rigorously guarantee that the Nash equilibrium coincides with the socially optimal solution; and unify treatment across atomic aggregative games and nonatomic routing games. We prove that every fixed point corresponds to an optimal incentive and derive sufficient conditions for global convergence. Numerical validation confirms these conditions hold and convergence is rapid in both canonical game settings.
This paper studies optimal contract design in sequential exploration: an agent sequentially opens $n$ costly boxes—each containing a prize—and selects one, while the principal commits upfront to a nonnegative payment contract to incentivize the agent and maximize expected utility (prize value minus payment). It bridges contract theory and the Pandora’s Box problem. Methodologically, it integrates game-theoretic modeling, dynamic programming, optimal stopping theory, and probabilistic analysis. The contributions are threefold: (i) the first polynomial-time algorithm for computing the exact optimal linear contract; (ii) a closed-form characterization of the optimal general (nonlinear) contract under single-prize and i.i.d. box assumptions; and (iii) theoretical guarantees of both principal-optimal utility and incentive compatibility. The framework yields a computationally tractable and interpretable mechanism design paradigm for sequential decision-making under uncertainty.
This paper studies a multi-armed bandit (MAB) problem where strategic agents serve as “arms”: each agent can manipulate reported rewards and incurred costs, necessitating an incentive-compatible mechanism that elicits high-performance truthful behavior while ensuring robust performance under non-equilibrium behavior (e.g., irrationality or deviations). To this end, we propose the first MAB framework jointly guaranteeing incentive compatibility and non-equilibrium robustness. We identify a key structural property enabling synergistic achievement of both objectives and integrate insights from second-price auctions to handle settings with no prior knowledge of arm qualities. Theoretically, our algorithm yields a non-vacuous lower bound on cumulative reward under arbitrary agent behavior; moreover, even without knowledge of true arm performances, it achieves an $O(sqrt{T})$ regret upper bound—substantially improving upon conventional approaches that either ignore incentives or lack robustness guarantees.
This paper investigates collective learning failure among myopic agents—employing greedy, exploration-free policies—in the multi-armed bandit (MAB) framework for social learning. Under a sequential decision-making setting with no private signals—where agents rely solely on shared history of actions and rewards—we establish, for the first time, that moderately myopic greedy strategies (e.g., ε-greedy or UCB variants with confidence intervals) incur linear regret. We precisely characterize the phase-transition threshold between myopia severity and exploratory capacity. By integrating social learning dynamics modeling with refined regret analysis, we derive tight upper and lower bounds, revealing general conditions under which greedy algorithms systematically fail. A key theoretical contribution is proving that “moderate optimism”—formalized as appropriately calibrated upper-confidence bonuses—is both necessary and sufficient to restore logarithmic regret. This provides a rigorous foundation for designing distributed learning protocols endowed with provably effective active exploration.
This work addresses the challenge of learning product valuations from noisy feedback in repeated contextual procurement auctions, while simultaneously ensuring incentive compatibility and minimizing social welfare loss. The authors propose two novel mechanisms: an Explore-then-Commit mechanism and a Frozen-Payments UCB mechanism, both enabling a tunable trade-off between incentive violation and social welfare regret. The former achieves a regret bound of Õ((ng)^{1/3}T^{2/3}), while the latter allows flexible balancing—via parameter tuning—between Õ(√(ngT)) regret and Õ(T^{3/4}) incentive error, or achieving both at Õ(T^{2/3}). Combining techniques from multi-armed bandits, contextual modeling, and mechanism design, the paper establishes corresponding lower bounds, demonstrating the near-optimality of the proposed approaches.
This work addresses the inadequacy of traditional Stackelberg equilibria when facing a boundedly rational follower employing the Upper Confidence Bound (UCB) algorithm for experiential learning, as the follower’s optimistic exploration can be strategically exploited. The authors propose a two-stage deception mechanism: in an initial baiting phase, the leader artificially inflates the UCB index of a target action; subsequently, in a trapping phase, the leader switches to a self-interested policy, inducing the follower—relying on the manipulated history—to persistently select that action. This approach constitutes the first provably effective deceptive mechanism for leaders confronting UCB followers, revealing a fundamental incompatibility between static equilibrium concepts and dynamic learning incentives. Under standard separation and payoff assumptions, the leader’s cumulative utility strictly exceeds the classical Stackelberg equilibrium upper bound, with manipulation cost bounded within an $O(\sqrt{T \ln T})$ regret term.
This work proposes a novel architecture based on multi-scale context awareness and dynamic graph reasoning to address the limited representational capacity of existing methods in complex scenes. By effectively fusing local details with global semantic information and introducing a learnable graph structure to model inter-entity relationships, the proposed approach significantly enhances the model’s ability to capture fine-grained features. Extensive experiments demonstrate that the method achieves state-of-the-art performance across multiple benchmark datasets, exhibiting particularly strong robustness and generalization under challenging conditions such as occlusion and scale variation.
This work addresses the limitations of traditional Bayesian incentive compatibility, which fails when agents possess private information, lack a common prior, or operate in incomplete information environments. The paper proposes a more general incentive exploration framework that dispenses with Bayesian assumptions and requirements of complete information, allowing agents to act according to any undominated strategy. By introducing a novel definition of incentive compatibility that does not rely on a common prior and integrating multi-prior robust optimization with non-Bayesian decision theory, the framework effectively handles action ties. This approach substantially extends the applicability and robustness of incentive-compatible mechanisms under information asymmetry and uncertainty.
本文通过引入随机性改进了动态委托-代理问题中的激励设计,使用Heteroscedastic GP-UCB算法解决了因不确定性导致的计算难题。