Score
Designs, implements, and analyzes algorithms and solution methods for sequential decision and optimization problems formulated as dynamic programs — covering exact and approximate dynamic programming, DP on DAGs/trees/decompositions, dynamic optimization and pricing algorithms, and dynamic (including fully dynamic) submodular maximization. Works produce near‑optimal decision policies, approximation schemes, efficient DP implementations and reductions (e.g., to matrix operations), and analyses of computational and performance trade‑offs in evolving systems.
This study investigates the performance trade-offs of data-driven approaches in finite-horizon dynamic pricing, with a focus on complex settings involving high-dimensional multi-product offerings, heterogeneous demand structures, and intertemporal revenue constraints. By systematically comparing Fitted Dynamic Programming (Fitted DP) against several reinforcement learning algorithms—integrating demand estimation, trajectory sampling, and expectation-based optimization—the work comprehensively evaluates their relative strengths in terms of revenue generation, stability, constraint satisfaction, and computational scalability. The findings reveal that Fitted DP exhibits superior scalability in structured, complex environments, whereas reinforcement learning demonstrates greater flexibility and adaptability. These insights provide both theoretical grounding and empirical evidence to inform method selection for real-world dynamic pricing systems.
In column generation, pricing subproblems heavily rely on application-specific structural properties, resulting in low reusability of customized solvers. To address this, we propose a domain-agnostic, general-purpose pricing framework that— for the first time—integrates dynamic programming into both the column generation and branch-and-price processes, serving as a transferable pricing solver independent of problem-specific algorithms. Our method uniformly handles diverse combinatorial structures without requiring redesign of pricing logic for each problem class. Evaluated on seven canonical integer programming problems, it consistently outperforms state-of-the-art commercial and open-source solvers in solution quality and runtime, achieving superior scalability and robustness. The framework significantly enhances the generality, reliability, and computational efficiency of large-scale exact optimization.
The absence of a universal, domain-agnostic dynamic programming (DP) modeling paradigm hinders systematic DP application to combinatorial optimization. Method: This paper introduces Domain-Independent Dynamic Programming (DIDP), a paradigm that decouples problem modeling from solving. We design DyPDL—a formal language for specifying DP models—and develop CAASDy, a general-purpose solver that unifies classical DP, A* search, and cost-algebraic state-space search within a verifiable DP framework for the first time. CAASDy supports interoperable interfaces with MIP and CP models, enabling fair empirical comparisons. Results: Experiments across multiple standard combinatorial optimization benchmarks demonstrate that CAASDy significantly outperforms leading commercial MIP and CP solvers. These results validate DIDP’s triple innovation: modeling generality, solving efficacy, and theoretical verifiability.
This paper addresses a class of multistage stochastic programming problems characterized by continuous state and action spaces, decision-dependent uncertainty, and limited statistical learning capability. To overcome the expressive limitations of conventional models, we propose an extended policy graph framework that explicitly captures the feedback effect of decisions on uncertainty and incorporates online learning mechanisms. Building upon this, we design a novel stochastic dual dynamic programming (SDDP) algorithm and its nonconvex approximation variant, tailored for efficiently solving such structured Markov decision processes. Experimental results on a suite of benchmark instances—increasing in complexity—demonstrate that our approach significantly improves policy quality and computational scalability. The work establishes a new paradigm for stochastic optimization that jointly integrates statistical learning with sequential decision-making, offering both enhanced expressiveness and tractability.
This work studies a class of combinatorial online learning problems—such as optimal binary search trees (BSTs), matrix-chain multiplication, and knapsack—that admit efficient dynamic programming (DP) solutions. It focuses specifically on online BST optimization under dynamically varying access frequencies: in each round, a low-average-search-cost tree must be selected rapidly to minimize cumulative regret against the best fixed offline BST. The paper innovatively extends the Hedge and Component Hedge algorithms to DP-induced exponentially large combinatorial structures, introducing a novel weight-propagation mechanism based on subproblem decomposition, probabilistic tree construction, and a corresponding regret analysis framework. Theoretically, it achieves a sublinear regret bound—strictly superior to brute-force enumeration or greedy baselines. Moreover, the approach is generalizable, enabling natural online adaptation of diverse DP-based problems.
This work addresses monotone submodular optimization under dynamic knapsack constraints in a multi-task setting. The authors propose a multi-task evolutionary optimization framework that enables efficient knowledge transfer across tasks sharing the same submodular objective but subject to different constraints. By constructing a compact Pareto front within a unified cost structure, this study is the first to integrate multi-task Pareto optimization with dynamic-constrained submodular maximization. Theoretical analysis demonstrates that the proposed algorithm achieves a $(1 - 1/e)$-approximation guarantee for each task in expected polynomial time. Empirical evaluations on the maximum coverage problem show that the method significantly outperforms existing baselines across various budget configurations.
This work addresses the problem of efficiently and exactly solving discounted Markov decision processes (DMDPs) for the optimal value function and policy. We propose a novel reduction framework that decomposes the exact solution into two subproblems: policy evaluation and computation of an approximately optimal value function. Leveraging state-of-the-art techniques in approximate dynamic programming, we design both deterministic and randomized algorithms tailored to these subtasks. Our approach achieves significantly improved computational efficiency, yielding the fastest known exact DMDP solver to date. The resulting algorithms demonstrate clear advantages over existing methods, both theoretically—through tighter complexity bounds—and empirically—via superior practical performance.
This work addresses the challenge of coupled signaling and control in partially observable Markov decision processes (POMDPs) by proposing an information-theoretic meta-dynamic programming framework under average-cost constraints. By introducing a dual-coupled information state—comprising the posterior distribution of the system state and the distribution over this posterior as sufficient statistics—the optimal stochastic control policy is decomposed into a separated structure dependent solely on these two information states, and necessary and sufficient conditions for optimality are established. The approach integrates directed information measures, Bayesian recursions, and dynamic programming over the probability simplex, naturally reducing to the classical POMDP solution in the absence of communication. This framework provides the first analytically tractable optimal control solution for POMDPs with endogenous information constraints, laying a theoretical foundation for integrated communication-control intelligent systems.
This study investigates the preservation and transfer of optimality properties in dynamic programming problems under diverse preference structures. By introducing conjugacy theory from dynamical systems and order-isomorphism techniques, it establishes, for the first time, rigorous equivalence relations among Epstein–Zin preferences, multiplicative Kreps–Porteus preferences, and risk-sensitive preferences, thereby enabling the cross-model transfer of optimality characteristics. This unified framework not only provides a coherent characterization of optimality across multiple preference specifications but also significantly enhances computational accuracy in applied settings: when implemented in a multisector real business cycle model, it improves the numerical precision of value function approximation by up to two orders of magnitude.
This work addresses the fully dynamic submodular maximization problem, which supports both insertions and deletions, by proposing the first general algorithmic framework that achieves constant-factor approximation guarantees with sublinear adaptive complexity. Built upon a dynamic streaming model, the framework integrates greedy strategies with a buffering mechanism and applies to both cardinality and rank-$k$ matroid constraints. Specifically, under a cardinality constraint, it attains a $(1/2 - O(\varepsilon))$-approximation with $O(1/\varepsilon^2)$ adaptivity; under a rank-$k$ matroid constraint, it achieves a $(1/4 - O(\varepsilon))$-approximation with $O(\log k / \varepsilon^2)$ adaptivity. This represents a significant advance over prior approaches, which were limited to insertion-only settings.