🤖 AI Summary
This study addresses the precise characterization of optimal policies and regret bounds in high-dimensional Bayesian linear bandits when the time horizon scales proportionally with the dimension. By leveraging Gaussian posterior identities and L1 convergence theory, the authors derive the limit of posterior uncertainty. They determine parameter overlap without requiring adaptive recursive closure assumptions and establish the optimality of greedy selection based on the posterior mean. The primary contributions include deriving a closed-form cumulative regret curve that decouples information acquisition from decision quality, and identifying asymptotically optimal policies. Furthermore, this work quantifies the additional regret ratio of Thompson sampling relative to greedy selection, which lies between one and two, thereby providing a rigorous theoretical benchmark for high-dimensional online learning.
📝 Abstract
We study Bayesian linear bandits with an isotropic Gaussian parameter, independent Gaussian candidate arms, and Gaussian reward noise when the horizon is proportional to the dimension. The normalized posterior uncertainty has an explicit limit that is uniform over all causal policies. Gaussian posterior identities then determine the limiting parameter overlaps without an assumed closure of the adaptive recursion. These results yield exact regret curves for Thompson sampling, posterior-mean greedy selection, and a family of policies that scale the posterior sampling covariance. The normalized realized cumulative regret converges in L1, uniformly on compact proportional-time intervals. A policy-uniform lower bound identifies the limiting optimal Bayes regret and proves that posterior-mean greedy selection attains it. Thompson sampling incurs a strictly larger leading regret; its instantaneous regret ratio relative to greedy selection lies between one and two and approaches two at long proportional horizons. Closed-form cumulative curves also identify a different comparison in the vanishing-noise limit. Finally, the instantaneous regret converges to a nondegenerate Gaussian decision-loss distribution, rather than to its mean. The analysis separates the amount of information acquired by a bandit policy from the quality of the decisions made using that information.