Score
Design and analyze online bandit algorithms whose exploration and update rules adapt to different feedback regimes—ranging from semi‑bandit and two‑point feedback to multi‑point or extra‑feedback settings—so they exploit additional observations to tighten regret bounds and obtain best‑of‑both‑worlds guarantees.
This work investigates a partial feedback setting in online learning that lies between full information and the multi-armed bandit, where selecting an action reveals not only its own loss but also the losses of additional actions. The paper proposes two novel algorithms based on implicit exploration: the first achieves near-optimal regret guarantees in general settings without requiring prior knowledge of the observation structure, while the second is tailored to combinatorial optimization problems, simultaneously ensuring computational efficiency and strong theoretical performance. Both algorithms improve upon existing approaches in both information-theoretic and computational terms, and they are the first to attain near-optimal regret in environments with unknown observation structures.
This paper addresses the long-standing open problem in online learning of simultaneously achieving constant regret against a given benchmark policy and √T regret against the hindsight-optimal policy. Focusing on symmetric zero-sum games—both normal-form and extensive-form—it unifies no-regret learning with exploitability. The authors propose the first bandit-feedback algorithm that provably attains both objectives optimally: O(1) regret relative to a prescribed baseline policy and O(√T) regret relative to the best-in-hindsight policy. The algorithm integrates game-theoretic analysis, online convex optimization, and bandit mechanisms. It achieves an optimal trade-off between robustness—suffering at most O(1) loss against adversarial opponents—and adaptivity—gaining Ω(T) reward against exploitable opponents—thereby breaking the performance limitations of conventional no-regret algorithms or minimax strategies.
This paper addresses the cascaded inference decision problem in edge intelligence scenarios, where the objective is to dynamically balance model accuracy and error probability across multiple models to minimize cumulative regret. We formulate each “arm” as an inference model characterized by its accuracy and error probability, embedded within a cascaded feedback structure. To this end, we propose an adaptive online learning–based decision framework. Theoretically, we prove that both the Lower Confidence Bound (LCB) and Thompson Sampling strategies achieve *O*(1) constant regret—substantially outperforming static or phased strategies such as Explore-then-Commit and Action Elimination. Our analysis further reveals that adaptive confidence updating is critical for overcoming the limitations of fixed-order execution. Extensive simulations validate the framework’s efficiency and robustness under uncertain edge environments.
This paper addresses the problem of safely and efficiently improving online learning performance for stochastic multi-armed bandits (MAB) and combinatorial MAB under offline–online distribution shift. To tackle distribution mismatch between offline and online data, we propose MIN-UCB—a novel adaptive algorithm that automatically discards low-quality offline data when the shift magnitude is unknown, ensuring safety, and actively leverages bounded-shift offline information to substantially reduce regret when the shift is bounded. We establish, for the first time, a theoretical lower bound on regret for MAB with biased offline data. We prove that MIN-UCB achieves tight instance-dependent and instance-independent regret bounds—strictly outperforming classical UCB. Both theoretical analysis and numerical experiments demonstrate its robustness and superiority across diverse shift regimes.
This paper investigates collective learning failure among myopic agents—employing greedy, exploration-free policies—in the multi-armed bandit (MAB) framework for social learning. Under a sequential decision-making setting with no private signals—where agents rely solely on shared history of actions and rewards—we establish, for the first time, that moderately myopic greedy strategies (e.g., ε-greedy or UCB variants with confidence intervals) incur linear regret. We precisely characterize the phase-transition threshold between myopia severity and exploratory capacity. By integrating social learning dynamics modeling with refined regret analysis, we derive tight upper and lower bounds, revealing general conditions under which greedy algorithms systematically fail. A key theoretical contribution is proving that “moderate optimism”—formalized as appropriately calibrated upper-confidence bonuses—is both necessary and sufficient to restore logarithmic regret. This provides a rigorous foundation for designing distributed learning protocols endowed with provably effective active exploration.
This work addresses the challenge of integrating offline data with online learning in stochastic linear bandits by proposing a novel algorithm that initially leverages offline data to form a prior and subsequently enhances online exploration in an adaptive manner, dynamically balancing the contributions of both sources. Within the structured linear bandit setting, the method is the first to simultaneously outperform purely online and purely offline strategies, achieving a sublinear regret bound with respect to the optimal action. Notably, the regret decreases as the amount of offline data increases. Theoretical analysis establishes a rigorous upper bound on regret, and extensive experiments demonstrate that the proposed approach significantly outperforms existing baselines across various settings.
This work addresses online convex optimization in adversarial environments under bandit feedback, where only the loss values at two queried points are observable. Focusing on μ-strongly convex loss functions, the paper proposes a novel algorithm based on two-point gradient estimation and introduces high-probability analysis techniques tailored to handle heavy-tailed noise, thereby overcoming the limitations of conventional concentration inequalities. The authors establish, for the first time, a high-probability regret bound of O(d(log T + log(1/δ))/μ), which is minimax optimal in both the time horizon T and the dimension d. This result resolves a long-standing open problem in the field and represents a significant advance in the theory of bandit online learning.
This work addresses constrained convex optimization with imperfect gradient predictions, where even with accurate predictions, a regret lower bound of Ω(√T) persists under single-point feedback. The paper proposes the TP-VR-OPT algorithm, which, under two-point feedback, establishes—for the first time—an information-theoretic regret lower bound of Ω(√𝔼[S_T]) that depends on the expected cumulative prediction error 𝔼[S_T], and achieves a matching upper bound of O(√(d𝔼[S_T])), differing only by a √d factor. The method integrates variance-reduced gradient estimation, optimistic online learning, and adaptive step sizes, requiring no prior knowledge of either S_T or the time horizon T. Furthermore, it naturally extends to non-stationary environments, maintaining adaptivity with respect to both dynamic path length and prediction error.
This work addresses the problem of agents actively refining their own features during online learning to obtain more favorable labels. It introduces an enhanced agent learning framework that, for the first time, supports multiclass classification, incorporates arm feedback, and respects budget constraints. By integrating feature refinement costs and a budget mechanism, the paper formulates an online learning model that embeds strategic agent behavior as a game-theoretic component and defines a novel combinatorial dimension to characterize learnability within this setting. Theoretical analysis establishes necessary and sufficient conditions for learnability under this paradigm, thereby extending the boundaries of online learning theory and providing a rigorous foundation for real-world systems that must balance agent incentives with learner performance.