🤖 AI Summary
This study addresses the challenge of nonlinear reward modeling in contextual bandits, where achieving both efficient online updates and low memory consumption remains difficult. To this end, this work proposes a multi-level discretized representation framework based on residual quantization. By mapping continuous contexts to discrete centroids via hierarchical codebooks and integrating a dynamic shadowing mechanism with additive bandit algorithms, the proposed method transcends the expressiveness limitations of linear models while preserving their O(1) update efficiency, thereby enabling nonlinear modeling under strictly bounded memory. Experimental results across 13 datasets demonstrate that this approach consistently outperforms existing baselines, matching the performance of XGBoost and neural networks while requiring only one-thousandth of their memory footprint.
📝 Abstract
Contextual bandits require balancing nonlinear reward modeling with online efficiency. Tree ensembles and neural methods capture nonlinearities but require periodic retraining and large replay buffers. Linear models update efficiently per observation with O(1) memory, but are fundamentally restricted to linear reward structures. We propose Residual Quantization (RQ) as a representation layer to bridge this gap. An offline-trained RQ codebook maps continuous contexts into discrete centroid assignments across multiple levels, set dynamically through a shadow mechanism. This enables a spectrum of additive bandit algorithms that achieve nonlinear expressivity with strictly bounded memory. Across 13 datasets, RQ variants beat their non-RQ counterparts on 11 of 13 datasets, often by wide margins, while matching doubling-retrain XGBoost and neural baselines using up to 1000 times less memory.