🤖 AI Summary
This study addresses the problem of learning an unknown objective function in online inverse linear optimization. It proposes a deterministic polynomial-time algorithm that builds upon a variant of the variable metric method, innovatively introducing a query-point distance trigger mechanism to dynamically revoke metric updates. This approach effectively overcomes the high complexity bottleneck inherent in prior randomized algorithms. The work establishes for the first time that optimal regret bounds can be achieved in polynomial time: it attains an optimal regret of O(√d) at any time step, with running time growing polynomially in both the dimension d and the horizon T. Consequently, this research successfully unifies theoretical optimality with computational efficiency.
📝 Abstract
In online inverse linear optimization, a learner recommends an action and then observes the choice of an expert who maximizes a fixed, unknown linear objective on $\mathbb{R}^{d}$; the goal is to learn to optimize this objective without observing it. Sakaue recently obtained the optimal regret $O(\sqrt d)$ with a randomized algorithm making $(dT)^{O(d)}$ linear optimizations per round, and asked whether it can be attained in polynomial time. We answer positively: our deterministic algorithm has regret $O(\sqrt d)$ for every horizon $T$ and runs in time polynomial in $d$ and $T$. It is a variant of the variable-metric algorithms of Sakaue et al.\ and Cai et al., in which a metric update is revoked once the query point moves far enough from where the update was made.