Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过多尺度矩阵权重解决在线逆线性优化问题,提出一种随机算法,在未知时间范围情况下达到O(√d)的期望遗憾界。
📝 Abstract
We study online inverse linear optimization with a fixed unknown linear utility: in each round, an environment presents a compact action set, the learner recommends an action from it, and the environment returns an action that maximizes the utility over the same set. When the utility vector and the actions lie in the $d$-dimensional Euclidean unit ball, we give a randomized algorithm whose regret---the cumulative utility shortfall relative to optimal actions---is $O(\sqrt d)$ in expectation for every time horizon, without knowledge of the horizon. The dependence on $d$ is optimal up to a constant factor by the known $Ω(\sqrt d)$ lower bound for horizons $T\ge d$. Our algorithm maintains matrix multiplicative weights on polynomial feature spaces at geometrically spaced scales. It selects a recommendation distribution by solving a linear program and updates its score matrices by comparing the available actions with the feedback action. With rational oracle outputs and feedback actions, an implementation computable relative to a linear-optimization oracle preserves the $O(\sqrt d)$ regret bound. Whether the same rate is attainable with running time polynomial in the dimension, horizon, and input length remains open.
Problem

Research questions and friction points this paper is trying to address.

online inverse linear optimization
regret bound
unknown linear utility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Online Inverse Linear Optimization
Multiscale Matrix Weights
Regret Bound
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.