misspecification reduction

Designs and analyzes algorithmic reductions that transform nonstationary or misspecified online decision problems into sequences of simpler misspecified or stationary bandit problems by partitioning the time horizon into blocks, upper-bounding within-block parameter drift, and converting dynamic-regret objectives into stationary or misspecified-bandit regret bounds. Builds blockwise decomposition schemes and accompanying proof bounds that quantify how misspecification error accumulates and enable reuse of existing misspecified-bandit guarantees.

misspecificationreduction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of non-stationary linear bandits with time-varying action sets and drifting reward parameters, unifying the treatment of general compact decision sets and K-armed contextual settings without requiring orthogonal structure assumptions. By partitioning the time horizon into blocks, the dynamic environment is transformed into a sequence of static linear bandit subproblems subject to parameter misspecification. The authors introduce a misspecification-sensitive restarting strategy that, for the first time, achieves the optimal dynamic regret bound of $\tilde{O}(T^{2/3}P_T^{1/3})$ simultaneously in both settings. This approach leverages a novel perspective based on misspecification reduction, effectively integrating techniques from time discretization, modeling, and dynamic regret analysis, thereby overcoming the structural assumptions prevalent in existing theoretical frameworks.

dynamic regretmisspecificationnon-stationary linear bandits

This work addresses the challenge in stochastic non-stationary linear bandits where time-varying parameters cause conventional methods to prematurely discard historical data, thereby losing valuable information. To mitigate this issue, the authors propose decomposing the reward model into stationary and non-stationary components and introduce invariance modeling—leveraging stable structures within historical data to reduce the effective problem dimensionality—within a dynamic regret minimization framework. The proposed ISD-linUCB algorithm integrates contextual linear modeling, dynamic weighting, and invariance identification to enable efficient online decision-making in non-stationary environments. Both theoretical analysis and empirical experiments demonstrate that, particularly in rapidly changing settings with sufficient historical data, the method achieves significantly lower dynamic regret compared to existing approaches.

dynamic regrethistorical datainvariance

Tracking Most Significant Shifts in Infinite-Armed Bandits

Jan 31, 2025
JS
Joe Suk
🏛️ Columbia University | Seoul National University

This paper studies the infinite-armed nonstationary multi-armed bandit (MAB) problem, where arm mean rewards drift over time in an unknown, arbitrary manner and no prior knowledge of nonstationarity parameters is available. To address this challenge, we propose a fully parameter-free online learning framework: first, a statistically significant reward drift detection mechanism enables adaptive identification of critical drifts; second, a black-box transformation coupled with randomized elimination adapts classical finite-arm algorithms to the infinite-arm setting; third, under mild reservoir distribution assumptions, we establish, for the first time, a minimax-optimal regret bound and uncover the distinctive phenomenon that “increasing rewards do not exacerbate learning difficulty.” Theoretical analysis yields a tight, adaptive regret rate depending solely on rotting-type nonstationarity—significantly improving upon existing parameter-dependent approaches.

Infinite ChoiceMulti-armed Bandit ProblemTime-varying Probabilities

Adaptive Smooth Non-Stationary Bandits

Jul 11, 2024
JS
Joe Suk
🏛️ Columbia University

This paper studies the non-stationary $K$-armed bandit problem with Hölder-continuous reward functions—characterized by smoothness exponent $eta$ and coefficient $lambda$—and time-varying rewards, focusing on dynamic regret minimization. We propose the first fully adaptive algorithm that requires no prior knowledge of $eta$ or $lambda$, integrating sliding-window estimation, adaptive segmentation, and a novel “safe-arm” mechanism. Our theoretical contributions are threefold: (1) We establish a tight minimax lower bound $Thetaig(T^{(1+eta)/(1+2eta)} lambda^{1/(1+2eta)}ig)$ for dynamic regret in general Hölder non-stationary environments; (2) We achieve full-range adaptivity—optimal for all $eta > 0$—for the first time; (3) We show that the safe-arm mechanism enables gap-dependent dynamic regret $O(log T)$, breaking the conventional $Omega(sqrt{T})$ barrier. These results resolve the long-standing open problem of adaptive dynamic regret control.

Analyzes K-armed non-stationary bandits with smooth reward changes.Establishes minimax dynamic regret rates for all parameters.Explores gap-dependent regret rates in non-stationary bandit environments.

Nonstationary Bandit Learning via Predictive Sampling

May 04, 2022
YL
Yueyang Liu
🏛️ Stanford University

This paper addresses the failure of Thompson sampling to maintain effective exploration in non-stationary multi-armed bandits due to its neglect of information timeliness. We propose Predictive Sampling, the first method to explicitly incorporate timeliness modeling into the Bayesian decision framework. It achieves adaptive exploration prioritization via dynamic prior updating and timeliness-weighted sampling. Theoretically, we derive the first Bayesian regret upper bound applicable to non-stationary environments and prove its boundedness. Computationally, we design a scalable approximate posterior inference mechanism. Experiments across diverse non-stationary settings—including abrupt and gradual distributional shifts—demonstrate that Predictive Sampling significantly outperforms classical Thompson sampling. The algorithm exhibits strong convergence properties, robustness to environmental dynamics, and practical scalability, making it suitable for real-world deployment.

Addresses poor Thompson sampling performance in non-stationary bandit environmentsProposes predictive sampling to prioritize long-term useful informationProvides scalable algorithms with theoretical guarantees for non-stationary bandits

Latest Papers

What's happening recently
View more

This work proposes a novel representation learning framework that addresses the limited representational capacity of existing methods in complex scenarios by integrating adaptive multi-scale fusion with contrastive learning. The approach dynamically aggregates multi-level semantic information and incorporates a structure-aware contrastive loss to enhance the model’s ability to discriminate fine-grained differences. Experimental results demonstrate that the proposed framework consistently outperforms state-of-the-art methods across multiple benchmark datasets, exhibiting particularly strong robustness under low-resource settings and in the presence of noise. These improvements yield higher-quality feature representations that significantly benefit downstream tasks.

change detectionMarkov chainspiecewise-stationary

This work addresses the long-standing open problem of minimizing dynamic regret against arbitrary sequences of dynamic comparators in unconstrained adversarial linear bandits when only pointwise loss feedback is available. The paper proposes an adaptive algorithmic framework that does not require prior knowledge of the number of comparator switches, achieving—for the first time in linear bandits—an optimal dynamic regret bound with respect to any number of switches $S_T$. Building upon adaptive ensembling techniques from multi-armed bandits and incorporating a parameter-free design alongside refined analysis of dynamic comparators, the method attains a dynamic regret upper bound of $\mathcal{O}(\sqrt{d(1+S_T)T})$, up to logarithmic factors. This result significantly advances the theory of online learning in non-stationary environments.

adversarialdynamic regretlinear bandits

This work addresses the challenge of online recommendation under heterogeneous user preferences, non-stationary context distributions, and the requirement to consistently outperform a baseline policy. The problem is formulated as a linear contextual multi-armed bandit with non-stationary heteroscedastic noise. We propose the first algorithm that simultaneously handles preference heterogeneity, context drift, and baseline constraints by extending the MED strategy to the linear setting, incorporating variance-aware suboptimality gap estimation and a constraint violation control mechanism. Theoretical analysis establishes an instance-dependent regret bound of Õ(κ/Δ̃·d²·log T) and an expected number of constraint violations bounded by Õ(d). Empirical results demonstrate that the proposed method significantly outperforms conservative baselines that ignore either context drift or preference heterogeneity.

conservative constraintcontext driftcontextual bandits

This work addresses the lack of theoretical guarantees for KL-regularized contextual bandits and episodic reinforcement learning with function approximation under model misspecification. The authors propose a novel framework that introduces a KL-based misspecification notion to formally characterize non-realizability, integrating regression-oracle-based algorithms with Gibbs policy updates for online learning. Building on this framework, they establish a high-probability KL-regret bound that explicitly accounts for model misspecification. This bound not only recovers the standard realizable case as a special instance but also provides the first rigorous theoretical guarantee for algorithms operating under imperfect modeling assumptions. Consequently, the approach significantly enhances the applicability and robustness of KL-regularized methods in function approximation settings.

contextual banditsfunction approximationKL-regularized

This work investigates static and dynamic regret minimization for unconstrained linear bandits (uBLO) in adversarial environments. By introducing a perturbation mechanism, the uBLO problem is reduced to a standard online linear optimization (OLO) setting, enabling the integration of comparator-adaptive OLO algorithms. The proposed approach achieves, for the first time without any prior knowledge, high-probability optimal guarantees for both static and dynamic regret: the dynamic regret scales as $\tilde{O}(\sqrt{P_T})$ with respect to the path-length $P_T$, and an $\Omega(\sqrt{dT})$ lower bound is established for adversarial linear bandits over the unit Euclidean ball. This study presents the first high-probability bounds simultaneously covering static and dynamic regret, while also improving upon existing expected regret analyses.

adversarial banditsdynamic regretonline linear optimization

Hot Scholars

CF

Charles F. Manski

Northwestern University
econometrics and statisticsjudgment and decisionpublic policy
AV

Aki Vehtari

Professor, Aalto University
Bayesian analysisBayesian statisticsGaussian processesProbabilistic programming
BW

Benjie Wang

University of California, Los Angeles
Machine LearningArtificial IntelligenceCausal InferenceTractable Probabilistic Models