๐ค AI Summary
This study addresses the challenge of jointly optimizing order preparation and delivery under dynamic stochastic order arrivals. The authors propose a novel policy decomposition paradigm that models the fulfillment process as a Markov decision process with synchronization constraints, uniquely treating the preparation phase as a state-constrained filter to decouple the two stages. Through policy-level decomposition, the problem is split into a primary delivery scheduling problem and a preparation compatibility subproblem, which are iteratively refined via a feedback mechanism. The resulting DDF-VFA solution framework integrates large neighborhood search, neural network-based value function approximation, and rolling horizon optimization. Evaluated on real-world datasets, the approach significantly outperforms existing baselines, achieving lower fulfillment costs and demonstrating strong scalability.
๐ Abstract
Modern supply chains span diverse operational environments, ranging from e-commerce distribution networks to customized production-to-order manufacturing lines. Across these settings, operational efficiency depends on coordinating two highly interdependent stages: order preparation and downstream delivery. Although these stages are traditionally managed in isolation, real-world fulfillment systems must satisfy stringent delivery expectations under dynamic stochastic order arrivals. To bridge this gap, we introduce the Dynamic Order Fulfillment Problem (DOFP), a new problem class unifying logistical challenges previously studied separately. We model DOFP as a Markov decision process whose state and decision spaces are partitioned into preparation and delivery sub-spaces, linked by synchronization constraints. While recent approaches attempt to optimize both fulfillment stages simultaneously over myopic rolling horizons, our framework isolates and optimizes the downstream delivery policy, treating preparation strictly as a state-level constraint filter. To solve this, we develop the Decomposition-Driven Framework with Value Function Approximation (DDF-VFA), which utilizes a novel policy-level decomposition. This design partitions the search into a delivery-stage master problem and a preparation-stage compatibility subproblem, iteratively refined via feedback loops. DDF-VFA executes this strategy by combining a large-neighborhood search over partial delivery decisions with a neural-network value function approximation for the cost-to-go. Numerical illustrations on two example variants using real-world datasets show that DDF-VFA consistently outperforms benchmarks that optimize the two stages independently or jointly without decomposition. Finally, the framework naturally scales to accommodate additional real-world complexities such as batched or multi-stage preparation.