Score
Designs, implements, and analyzes posterior‑sampling (Thompson Sampling) algorithms for sequential decision problems that replace signed exploration noise with its absolute value and include single‑run and ensemble variants (ATS, EATS, ensemble TS with absolute noise). Work covers building these TS variants, ensuring optimism in expectation, preserving the computational efficiency of standard Thompson Sampling, and deriving frequentist-style regret guarantees for the modified algorithms.
This work proposes a novel architecture based on adaptive feature fusion and contrastive learning to address the limited generalization of existing methods in complex scenarios. By dynamically integrating multi-scale semantic information and incorporating a task-aware contrastive loss function, the model achieves enhanced robustness under cross-domain and few-shot settings. Extensive experiments demonstrate that the proposed approach significantly outperforms state-of-the-art models across multiple benchmark datasets, yielding an average accuracy improvement of 3.2%. Moreover, the method offers superior interpretability and computational efficiency, establishing a promising new direction for few-shot visual recognition.
This work addresses finite-horizon Markov decision processes (MDPs) with unknown rewards and state transitions, exhibiting complex temporal dependencies. We establish the first no-regret guarantee for Thompson sampling (TS) under a joint Gaussian process (GP) prior over rewards and transitions. The core analytical challenges stem from the non-Gaussianity of value functions and the intractability of Bayesian updates under multi-step Bellman recursion. To overcome these, we: (1) propose a joint GP model for rewards and transitions, and design a TS algorithm tailored to multi-step optimization; (2) develop the first analytical pathway for deriving no-regret bounds for TS within GP-based Bellman recursion; and (3) extend the elliptical potential lemma to the multi-output setting, thereby resolving the non-Gaussian value function bottleneck. Our analysis yields a regret bound of $ ilde{O}(sqrt{KHGamma(KH)})$, where $Gamma$ quantifies GP complexity. This result provides a foundational theoretical guarantee for structured Bayesian reinforcement learning.
To address the challenge of jointly maximizing predictive mean, uncertainty, and minimizing intra-batch redundancy in batch Bayesian optimization (BO), this paper proposes Thompson Sampling with Regret-to-Sigma Ratio (TS-RSR). TS-RSR is the first method to incorporate the regret-to-sigma ratio into the batch acquisition objective, leveraging a Thompson sampling approximation to explicitly balance exploration and exploitation across batch points. We provide theoretical guarantees of convergence. By integrating Gaussian process modeling with an efficient batch sampling scheme, TS-RSR significantly reduces intra-batch redundancy. Extensive experiments on synthetic benchmarks and real-world tasks demonstrate that TS-RSR consistently outperforms state-of-the-art batch BO methods, achieving new state-of-the-art performance.
This paper addresses sequential decision-making in non-stationary multi-armed bandits (NS-MAB) with time-varying, action-independent rewards, bridging a theoretical gap in Thompson sampling (TS) under non-stationarity. We propose two sliding-window TS algorithms: BETA-SWTS, based on Beta posterior inference, and γ-SWGTS, based on Gaussian posterior inference. Crucially, we develop the first unified regret analysis framework applicable to arbitrary non-stationary environments—including both abrupt changes and gradual drifts. Our framework introduces a novel error metric that rectifies fundamental limitations of prior analyses, enabling the first explicit, tight regret bounds for both algorithms across these two dynamic regimes. Theoretical analysis proves their superiority over classical approaches such as SW-UCB. Empirical evaluations confirm their robustness and faster convergence across diverse non-stationary benchmarks.
In stochastic linear multi-armed bandits with infinite action sets and finite ensemble sizes, existing methods suffer from high computational overhead due to ensemble scaling linearly with time horizon $T$. Method: This paper proposes a novel ensemble sampling framework that integrates Bayesian posterior approximation, linear function approximation, and concentration inequality analysis. Contribution/Results: Differing from conventional approaches requiring $O(T)$ base learners, our method is the first to achieve a lightweight ensemble of only $O(d log T)$ learners in structured bandits—breaking the linear dependence on $T$. We establish a regret upper bound of $ ilde{O}((d log T)^{5/2} sqrt{T})$ in $d$-dimensional linear environments over horizon $T$, approaching the optimal $ ilde{O}(sqrt{T})$ benchmark. The framework naturally accommodates infinite action spaces and significantly improves computational efficiency, offering a new paradigm for scalable, high-dimensional, long-horizon online decision-making under large or infinite action sets.
This work addresses the lack of high-probability regret bounds for Gaussian Process Thompson Sampling (GP-TS) in Bayesian optimization and the unclear dependence of the failure probability δ on the time horizon T. Under the assumption that the objective function is a sample path from a Gaussian process, the paper establishes the first polynomial-in-δ lower bound on the regret of GP-TS and derives an improved upper bound on cumulative regret. By relaxing conditions used in GP-UCB analyses and introducing refined probabilistic arguments alongside auxiliary lemmas, the authors provide several rigorous theoretical guarantees, including bounds on second moments and a relaxed expected regret bound. The results demonstrate that GP-TS exhibits superior dependence on both δ and T compared to prior analyses, thereby filling a critical theoretical gap and offering new foundations for Bayesian optimization.
This work addresses the inflexibility of traditional Bayesian sequential decision-making methods, which require full parametric modeling and struggle to incorporate structural constraints. The authors propose a minimalist Bayesian framework that places a prior only on the location of the optimal arm and leverages profile likelihood to eliminate nuisance parameters, yielding a generalized posterior that naturally accommodates structural assumptions such as unimodality. Building on this, they introduce the MINTS algorithm—the first to integrate minimalist Bayesian principles into Thompson sampling—enabling automatic adaptation to structural constraints without modeling irrelevant parameters while preserving theoretical optimality. In mean-constrained multi-armed bandits, MINTS achieves near-optimal non-asymptotic regret bounds; under no structure and unimodal settings, it attains the classic Lai–Robbins constant and a sharp asymptotic constant dependent solely on the neighborhood of the optimal arm, respectively.
This work addresses the challenge of Bayesian optimization when only pairwise preference feedback—rather than scalar rewards—is available. It proposes a novel Thompson sampling-based approach that models the latent utility difference via a monotonic link function and introduces a dueling kernel induced by base kernels to handle preference comparisons. The key contributions include the first finite-time regret bound for Thompson sampling under preference feedback, a dual-Thompson-sampling (dual-TS) pairing mechanism, and an anchor-invariance analytical framework. Theoretically, the method is shown to achieve convergence performance comparable to classical Thompson sampling with scalar rewards. Empirical evaluations on both synthetic and real-world datasets demonstrate its effectiveness and validate the theoretical claims.
This work addresses the limitation of conventional optimism-based approaches in stochastic generalized linear bandits, which struggle to yield variance-aware regret bounds. For the first time, the Gaussian Poincaré inequality is introduced into the analysis of Thompson sampling. By leveraging this functional inequality to control estimation errors after an initial warm-up phase, the method circumvents the need for optimistic confidence sets and establishes a regret upper bound that adapts to the reward variance. This result not only provides stronger theoretical support for the performance of Thompson sampling in generalized linear bandits but also opens a novel avenue for analyzing Bayesian bandit algorithms through functional inequalities.