Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits

๐Ÿ“… 2026-08-04
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the weighted generalized low-rank matrix problem under multiple objectives with lexicographic priorities by proposing the Lexi-LowGLM algorithm, which uniquely integrates lexicographic preference modeling with generalized low-rank matrix estimation. The method first estimates low-rank subspaces corresponding to each objective and then performs lexicographic online learning in the reduced feature space, employing Online Newton Step for efficient parameter updates. Theoretical analysis establishes a regret bound that depends only on the effective low-rank dimension $(d_1 + d_2)r$ rather than the ambient dimension $d_1 d_2$, while reducing computational complexity from $O(T^2)$ to $O(T)$. Empirical results demonstrate the algorithmโ€™s superior efficiency and performance.
๐Ÿ“ Abstract
This paper studies generalized low-rank matrix bandits with multiple prioritized objectives. At each round, the learner selects a matrix-valued arm and observes a vector-valued reward, whose components correspond to multiple objectives with different priority levels. Each objective is governed by an objective-specific generalized low-rank matrix model, and the learner evaluates arms according to a lexicographic preference order, prioritizing higher-level objectives before lower-level ones. We propose \textsc{Lexi-LowGLM}, an efficient online algorithm that first estimates objective-specific low-rank subspaces and then performs lexicographic learning in the reduced feature spaces. Unlike existing single-objective algorithms that repeatedly solve a batch generalized linear estimator using all historical observations, \textsc{Lexi-LowGLM} updates each objective-specific estimator via an online Newton step, reducing the estimator-update complexity over $T$ rounds from $O(T^2)$ to $O(T)$. We establish a regret bound of $\widetilde O\left(W_i^{\rm lex}\sqrt{m}\,(d_1+d_2)r\sqrt{T}\right)$ for each objective $i\in[m]$, where $r$ is an upper bound on the ranks of the objective-specific parameter matrices and $W_i^{\rm lex}$ characterizes the lexicographic trade-off effect. This bound depends on the effective low-rank dimension $(d_1+d_2)r$ rather than the ambient dimension $d_1d_2$. Numerical experiments further validate the effectiveness and computational efficiency of the proposed method.
Problem

Research questions and friction points this paper is trying to address.

generalized low-rank matrix bandits
multi-objective optimization
lexicographic preference
online learning
matrix-valued arms
Innovation

Methods, ideas, or system contributions that make the work stand out.

lexicographic bandits
generalized low-rank matrix
online Newton step
multi-objective reinforcement learning
dimensionality reduction
๐Ÿ”Ž Similar Papers
No similar papers found.