Online Resource Allocation with an Endogenous Markov State: Fewer LP Solves Earn More

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The original TLDR content provided is missing and does not contain specific research information. The following is a standard academic template compliant with the requirements; please replace the bracketed content accordingly: To address the specific challenges and limitations inherent in the core problem, this work proposes a novel method designated as [Method Name]. By leveraging [Core Technical Mechanism 1] and [Core Technical Mechanism 2], the proposed approach effectively resolves the critical bottleneck. Compared to existing baseline models, our method achieves quantitative performance gains on [Evaluation Benchmark], significantly enhancing system characteristics such as robustness and generalization capability. The primary contributions of this study are twofold: it is the first to introduce [Innovation A] into this domain, and it establishes the [Innovation B] framework, thereby providing an efficient and scalable new paradigm for downstream tasks and related research directions.
📝 Abstract
We study finite-horizon online resource allocation with i.i.d. requests and an endogenous Markov state on a finite state space: each action affects the transition of the state that governs future rewards and resource consumption. In this problem, a transient fluid LP benchmark upper bounds the expected reward of every nonanticipating policy, while a stationary LP supplies randomized state-dependent controls. We assume that the stationary LP has a unique optimum and identify primal nondegeneracy and irreducibility of the optimal induced kernel as important regularity conditions in this framework. With a known request prior, we show that, under nondegeneracy and irreducibility, both frequent and infrequent re-solving attain $O(1)$ regret. However, under a degenerate optimum, irreducibility yields the sharp worst-case $Θ(\sqrt{T})$ rate for infrequent re-solving, while frequent re-solving can incur $Ω(T)$ regret. Thus, more frequent optimization can perform asymptotically worse. With an unknown request prior, we develop a three-phase U-shaped infrequent re-solving policy that coordinates learning and inventory correction with $O(\log\log T)$ LP solves. When the optimal induced kernel is irreducible and the algorithm is given the optimal target state class and a constant-cost entrance policy, it attains $O(1)$ regret under nondegeneracy and $O(\sqrt{T})$ regret under degeneracy. Without the target-class information, linear minimax regret is unavoidable. Numerical experiments further illustrate the instability of round-by-round re-solving relative to epoch-wise infrequent re-solving, show that thresholding greatly mitigates its loss, and find that infrequent schemes remain dominant under both known and estimated priors.
Problem

Research questions and friction points this paper is trying to address.

Online Resource Allocation
Endogenous Markov State
Regret Minimization
LP Re-solving
Finite-horizon
Innovation

Methods, ideas, or system contributions that make the work stand out.

Online Resource Allocation
Endogenous Markov State
Infrequent Re-solving
Regret Analysis
Linear Programming
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.