Score
Designs and formalizes partially observable Markov decision process (POMDP) models in which actions are structured as a kind plus parameters (e.g., discrete action types with continuous or structured parameter vectors), specifying the state space, observation model, transition kernel, reward/cost function, and the parameterized action-space semantics. These models are built and analyzed to enable planning, policy synthesis, and decision reasoning under partial observability for sequential decision problems where action effects depend on both action type and its parameters.
This work proposes a method for learning the structure and parameters of discrete partially observable Markov decision processes (POMDPs) from action-observation sequences under weak assumptions—specifically, when the system state is unobservable and the action-induced transition matrices are rank-deficient. By integrating predictive state representations (PSRs) with tensor decomposition, the approach jointly estimates the observation and transition matrices via a similarity transformation and partitions states into equivalence classes sharing identical observation distributions, thereby constructing an explicit likelihood model. In contrast to conventional tensor-based methods that require full-rank actions and complete observability, this is the first approach to enable explicit modeling of both transition and observation likelihoods in partially observable settings. The learned model, given sufficient data, can be directly employed by standard POMDP solvers, achieving planning performance comparable to PSR-based methods while supporting behavior customization through explicit likelihoods.
This work addresses the problem of efficiently computing optimal values and policies in finite-horizon partially observable Markov decision processes (POMDPs) with multiple environment models, where the initial state is adversarially chosen. We first establish that this problem remains PSPACE-complete even in the multi-environment setting. Subsequently, we propose the first practical and computationally efficient algorithm that integrates explicit modeling of adversarial initial states with finite-horizon dynamic programming. Empirical evaluation on several standard benchmarks demonstrates that our approach significantly outperforms the only previously known algorithm, thereby confirming its effectiveness and practical utility.
Online planning for continuous POMDPs with high-dimensional continuous observations (e.g., images) suffers from prohibitive computational cost due to repeated evaluation of the original, expensive observation model, hindering real-time deployment. Method: We propose a planning framework based on a simplified observation model, integrating particle filtering, continuous POMDP solvers, and probabilistic bound decomposition. Crucially, we derive a provable performance lower bound using the total variation (TV) distance—enabling offline construction of the simplified model while eliminating all online queries to the original observation model. Contribution/Results: This is the first work to establish a TV-distance-based, theoretically guaranteed performance bound for simplified observation models in continuous POMDP planning. It ensures planning quality without runtime access to the original model and extends concentration theory for particle-belief MDPs. Empirical evaluation confirms seamless integration with existing online solvers and yields tight, computationally tractable performance bounds—even with zero runtime calls to the original observation model.
Discrete partially observable Markov decision processes (POMDPs) lack deterministic performance guarantees for online planning solutions. Method: This paper proposes an online planning framework that, for the first time, establishes a deterministic error bound between any time-bounded approximate solution and the optimal value function. The approach integrates POMDP modeling, online Monte Carlo tree search (MCTS), upper-confidence-bound propagation, and rigorous error bound derivation—designed as a plug-and-play enhancement to existing MCTS-based planners. Contribution/Results: Evaluated on standard benchmarks, the method achieves significantly improved solution quality with negligible increase in computational overhead, while providing verifiable theoretical guarantees. It is the first POMDP planning framework to simultaneously ensure real-time execution, deterministic error bounds, and plug-and-play compatibility with state-of-the-art tree search algorithms.
This paper addresses the foundational question: *When can a physical system be rigorously regarded as an agent possessing beliefs and goals?* We propose a POMDP-based explanatory framework, formalizing an agent as a physical system satisfying two joint constraints: (i) its state evolution must conform to Bayesian belief-updating dynamics, and (ii) its policy must be optimal with respect to a specified objective. Crucially, we introduce the *completeness of a POMDP solution*—simultaneous adherence to correct belief evolution and optimal action selection—as a necessary and empirically falsifiable criterion for agency, overcoming the limitations of prior definitions relying solely on belief mapping. This yields the first axiomatization of agency that is mathematically rigorous, computationally operational, and empirically testable. It reveals necessary constraints linking physical dynamics to agential properties, thereby establishing a theoretical foundation for AI safety verification and computational modeling of consciousness.
This study investigates the set of achievable value functions in infinite-horizon partially observable Markov decision processes (POMDPs) under memoryless stochastic policies. Addressing the long-standing lack of a precise mathematical characterization of this set, the work establishes for the first time that it forms a semialgebraic set. Specifically, it explicitly constructs a system of polynomial inequalities—derived from the system dynamics, observation kernel, and reward structure—that fully describes the feasible value functions, thereby revealing the nonlinear constraints and intricate geometric structure induced by partial observability. This result generalizes the classical finding that the value function set in fully observable MDPs is polyhedral, clarifies the dependence of achievable values on the initial state distribution, and uncovers novel phenomena such as isolated locally optimal policies, thus providing a new theoretical foundation for POMDP policy optimization.
This work addresses the optimal observability problem (OOP) in uncertain environments, which entails balancing task feasibility against sensing costs. Focusing on its decidable subproblems—sensor selection (SSP) and position observability (POP)—the paper proposes a novel solution framework based on POMDP decomposition, integrating parameter synthesis with a symbolic–subsymbolic hybrid approach. This method dramatically improves computational efficiency, scaling solvable instances by three orders of magnitude and reducing runtime by five orders of magnitude compared to prior techniques. Consequently, the approach substantially expands the tractable boundary of observability-aware planning in partially observable settings.
This work addresses the severe limitations imposed by the curse of dimensionality and the curse of history in partially observable Markov decision processes (POMDPs), which hinder policy performance under finite planning horizons. The paper introduces VOIMCP, an algorithm that, for the first time, dynamically integrates the value of information (VOI) into POMDP planning. Built upon a Monte Carlo tree search framework, VOIMCP evaluates the VOI of observations in belief space and selectively prunes low-value observation branches. This approach substantially improves computational efficiency while providing theoretical guarantees of near-optimality and non-asymptotic convergence bounds. Empirical results demonstrate that VOIMCP significantly outperforms existing baselines across multiple POMDP benchmarks, confirming its superior performance and efficiency under limited computational resources.
This work addresses the challenge of learning near-optimal policies in partially observable Markov decision processes (POMDPs) using only finite observation-action histories. The authors propose a hyper-state MDP framework that enables efficient model estimation from a single trajectory and computes near-optimal finite-window policies via value iteration. A key theoretical contribution is the novel connection established between filter stability and concentration inequalities for weakly dependent random variables, which yields tight sample complexity guarantees for single-trajectory estimation in the hyper-state MDP. By integrating model-based reinforcement learning, hyper-state modeling, and analysis of non-independent sequences, the approach rigorously approximates high-performance policies in the original POMDP while maintaining strong theoretical foundations.
This work addresses the lack of finite-time theoretical guarantees for Monte Carlo tree search (MCTS) in partially observable Markov decision processes (POMDPs) with continuous observation spaces. To this end, the authors propose Voro-POMCPOW, an algorithm that extends the UCB exploration mechanism and introduces an adaptive observation-space partitioning framework based on Voronoi cells. This approach effectively handles action-selection dependencies and non-stationarity while preserving the original observation generator and maintaining a finite branching factor. The paper provides the first finite-time theoretical analysis for MCTS in continuous POMDPs, establishing high-probability polynomial concentration bounds on root-node value estimates and finite-time bounds on partitioning error. Empirical results demonstrate that Voro-POMCPOW achieves competitive performance while offering strong theoretical guarantees and is readily extensible to continuous MDPs.