🤖 AI Summary
This study addresses the challenge of online learning with delayed feedback, where expert corrections are required but early accumulated errors can impede subsequent query decisions. To tackle this, we propose the ORUCB algorithm, which employs polynomial response pooling and confidence-weighted risk regression. By introducing a cumulative response learning error bound to calibrate risk, ORUCB effectively mitigates exploration bias arising from unfinished tasks, thereby achieving a principled balance between exploration and exploitation. Theoretically, we establish a pseudo-regret upper bound of $O(\sqrt{T}\log T)$. Empirically, experiments across four benchmark data streams demonstrate that our method consistently achieves significantly lower costs compared to seven baseline models.
📝 Abstract
An inaccurate expert can still provide useful information after correction. We study online learning to defer in which the learner chooses an expert and fixes a correction function before purchasing its answer, then applies that function to the answer received. The difficulty is that observed losses reflect both expert quality and an unfinished correction: early errors can discourage queries that would be valuable after learning. We propose ORUCB, which pools shared and expert-specific polynomial responses. A bound on cumulative response-learning error calibrates confidence-weighted risk regression and exploration, allowing the router to account for this error when deciding which answers to buy. Under bounded residuals and disagreements, a fixed feasible model of optimal responses, and linear models of free and optimal queried risk, the calibrated algorithm achieves high-probability pseudo-regret $O(\sqrt T\log(T+1))$ over $T$ rounds for fixed problem parameters. The guarantee permits singular answer distributions and misspecified shared responses; optimality is relative to the bounded response class. On four test streams, the selected cubic policy has lower fee-inclusive cost than seven baselines that deploy answers unchanged. Comparisons with a common correction learner examine routing, while six-price comparisons measure cost and query rates.