frozen-payment ucb

Designs, builds, or analyzes incentive-aware exploration mechanisms that combine UCB-style bandit selection with a payment-freezing rule: protocols that temporarily fix or suspend payments to agents to stabilize incentives while performing upper-confidence-bound exploration. Work includes specifying the UCB allocation and payment-freezing procedure, proving bounds on welfare regret and incentive error, and tuning freeze/learning parameters to target rates such as ~o(√(ngt)) or ~o(t^{2/3}).

frozen-paymentucb

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Principal-Agent Bandit Games with Self-Interested and Exploratory Learning Agents

Dec 20, 2024
JL
Junyan Liu
🏛️ University of Washington

This paper studies incentive mechanism design in principal-agent multi-armed bandit games, where the agent must autonomously explore an unknown environment—rather than being fully informed and greedy. Addressing the realistic setting where agents are self-interested and exploration is inherently uncertain, we propose the first robust elimination framework tailored to exploration-aware learning agents. Our method integrates adaptive search, robust incentive design, and a coupled Bayesian/frequentist estimation scheme with explicit exploration. Under i.i.d. rewards, it achieves a regret bound of $widetilde{O}(T^{2/3})$; under linear rewards, it attains the optimal $widetilde{O}(sqrt{T})$ bound—improving significantly over Dogan et al. (2023a)’s $widetilde{O}(T^{11/12})$. This provides the first theoretically optimal solution for principals to dynamically guide learning agents in exploration-exploitation settings.

Achieving improved regret bounds by enhancing robustness to agent exploration.Developing algorithms for i.i.d. and linear reward settings with bandit feedback.Modeling self-interested learning agents with exploration behaviors in principal-agent bandit games.

Adaptive Incentive Design with Learning Agents

May 26, 2024
CM
Chinmay Maheshwari
🏛️ UC Berkeley | Cornell University

Existing incentive mechanisms fail to adapt to learning agents whose strategies evolve continuously over time. Method: We propose a two-timescale adaptive incentive mechanism driven by individual externality—the difference between an agent’s marginal cost and the operator’s marginal cost—and updated at a rate slower than the agents’ learning dynamics. Leveraging two-timescale stochastic approximation and differential game theory, we develop a technical framework comprising externality modeling, fixed-point analysis, and convergence proof. Contribution/Results: We establish the first general incentive framework decoupled from agents’ learning dynamics; rigorously guarantee that the Nash equilibrium coincides with the socially optimal solution; and unify treatment across atomic aggregative games and nonatomic routing games. We prove that every fixed point corresponds to an optimal incentive and derive sufficient conditions for global convergence. Numerical validation confirms these conditions hold and convergence is rapid in both canonical game settings.

Aligning Nash equilibrium with socially optimal strategiesDesigning adaptive incentives for learning agents in gamesEnsuring convergence in atomic and non-atomic game settings

Designing Exploration Contracts

Mar 04, 2024
MH
Martin Hoefer
🏛️ RWTH Aachen University | Goethe University Frankfurt | University of Southern Denmark

This paper studies optimal contract design in sequential exploration: an agent sequentially opens $n$ costly boxes—each containing a prize—and selects one, while the principal commits upfront to a nonnegative payment contract to incentivize the agent and maximize expected utility (prize value minus payment). It bridges contract theory and the Pandora’s Box problem. Methodologically, it integrates game-theoretic modeling, dynamic programming, optimal stopping theory, and probabilistic analysis. The contributions are threefold: (i) the first polynomial-time algorithm for computing the exact optimal linear contract; (ii) a closed-form characterization of the optimal general (nonlinear) contract under single-prize and i.i.d. box assumptions; and (iii) theoretical guarantees of both principal-optimal utility and incentive compatibility. The framework yields a computationally tractable and interpretable mechanism design paradigm for sequential decision-making under uncertainty.

Computing optimal contracts for principal-agent scenarios with various assumptionsDesigning contracts to maximize principal's reward in exploration tasksOptimizing linear contracts for sequential search problems efficiently

Robust and Performance Incentivizing Algorithms for Multi-Armed Bandits with Strategic Agents

Dec 13, 2023
SA
Seyed A. Esmaeili
🏛️ Simons Laufer Mathematical Sciences Institute | University of Maryland | Microsoft Research

This paper studies a multi-armed bandit (MAB) problem where strategic agents serve as “arms”: each agent can manipulate reported rewards and incurred costs, necessitating an incentive-compatible mechanism that elicits high-performance truthful behavior while ensuring robust performance under non-equilibrium behavior (e.g., irrationality or deviations). To this end, we propose the first MAB framework jointly guaranteeing incentive compatibility and non-equilibrium robustness. We identify a key structural property enabling synergistic achievement of both objectives and integrate insights from second-price auctions to handle settings with no prior knowledge of arm qualities. Theoretically, our algorithm yields a non-vacuous lower bound on cumulative reward under arbitrary agent behavior; moreover, even without knowledge of true arm performances, it achieves an $O(sqrt{T})$ regret upper bound—substantially improving upon conventional approaches that either ignore incentives or lack robustness guarantees.

Design robust algorithms ensuring non-vacuous rewards under non-equilibrium behaviorHandle unknown arm performance using second-price auction-inspired methodsIncentivize strategic agents in multi-armed bandits to maximize performance

Bandit Social Learning: Exploration under Myopic Behavior

Feb 15, 2023
KB
Kiarash Banihashem
🏛️ University of Maryland | Microsoft

This paper investigates collective learning failure among myopic agents—employing greedy, exploration-free policies—in the multi-armed bandit (MAB) framework for social learning. Under a sequential decision-making setting with no private signals—where agents rely solely on shared history of actions and rewards—we establish, for the first time, that moderately myopic greedy strategies (e.g., ε-greedy or UCB variants with confidence intervals) incur linear regret. We precisely characterize the phase-transition threshold between myopia severity and exploratory capacity. By integrating social learning dynamics modeling with refined regret analysis, we derive tight upper and lower bounds, revealing general conditions under which greedy algorithms systematically fail. A key theoretical contribution is proving that “moderate optimism”—formalized as appropriately calibrated upper-confidence bonuses—is both necessary and sufficient to restore logarithmic regret. This provides a rigorous foundation for designing distributed learning protocols endowed with provably effective active exploration.

Analyzing learning failures in myopic bandit algorithmsProviding theoretical foundation for intentional exploration algorithmsStudying exploration under behavioral biases in social learning

Latest Papers

What's happening recently
View more

This work addresses the challenge of learning product valuations from noisy feedback in repeated contextual procurement auctions, while simultaneously ensuring incentive compatibility and minimizing social welfare loss. The authors propose two novel mechanisms: an Explore-then-Commit mechanism and a Frozen-Payments UCB mechanism, both enabling a tunable trade-off between incentive violation and social welfare regret. The former achieves a regret bound of Õ((ng)^{1/3}T^{2/3}), while the latter allows flexible balancing—via parameter tuning—between Õ(√(ngT)) regret and Õ(T^{3/4}) incentive error, or achieving both at Õ(T^{2/3}). Combining techniques from multi-armed bandits, contextual modeling, and mechanism design, the paper establishes corresponding lower bounds, demonstrating the near-optimality of the proposed approaches.

bandit learningcontextual procurement auctionsregret-incentive tradeoff

This work addresses the inadequacy of traditional Stackelberg equilibria when facing a boundedly rational follower employing the Upper Confidence Bound (UCB) algorithm for experiential learning, as the follower’s optimistic exploration can be strategically exploited. The authors propose a two-stage deception mechanism: in an initial baiting phase, the leader artificially inflates the UCB index of a target action; subsequently, in a trapping phase, the leader switches to a self-interested policy, inducing the follower—relying on the manipulated history—to persistently select that action. This approach constitutes the first provably effective deceptive mechanism for leaders confronting UCB followers, revealing a fundamental incompatibility between static equilibrium concepts and dynamic learning incentives. Under standard separation and payoff assumptions, the leader’s cumulative utility strictly exceeds the classical Stackelberg equilibrium upper bound, with manipulation cost bounded within an $O(\sqrt{T \ln T})$ regret term.

bounded rationalitydeceptionoptimism

This work proposes a novel architecture based on multi-scale context awareness and dynamic graph reasoning to address the limited representational capacity of existing methods in complex scenes. By effectively fusing local details with global semantic information and introducing a learnable graph structure to model inter-entity relationships, the proposed approach significantly enhances the model’s ability to capture fine-grained features. Extensive experiments demonstrate that the method achieves state-of-the-art performance across multiple benchmark datasets, exhibiting particularly strong robustness and generalization under challenging conditions such as occlusion and scale variation.

incentive designNash equilibriumnonlinear games

This work addresses the limitations of traditional Bayesian incentive compatibility, which fails when agents possess private information, lack a common prior, or operate in incomplete information environments. The paper proposes a more general incentive exploration framework that dispenses with Bayesian assumptions and requirements of complete information, allowing agents to act according to any undominated strategy. By introducing a novel definition of incentive compatibility that does not rely on a common prior and integrating multi-prior robust optimization with non-Bayesian decision theory, the framework effectively handles action ties. This approach substantially extends the applicability and robustness of incentive-compatible mechanisms under information asymmetry and uncertainty.

BayesianismCommon PriorFull-Information