🤖 AI Summary
This study addresses the absence of strong policies and the computational challenges associated with best-effort policy synthesis in fully observable non-deterministic (FOND) planning. To this end, it proposes a unified solving algorithm based on antichains. By leveraging set-theoretic minimal elements to compactly represent state sets, the method integrates both strong and best-effort planning into a single framework, yielding a complete and consistent planner capable of simultaneously producing strong solutions, weak solutions, and failure certificates. Experimental evaluations demonstrate that the proposed algorithm matches or exceeds mainstream tools in problem coverage, while achieving significantly superior runtime efficiency on large-scale instances compared to PR2 and BeSyftP.
📝 Abstract
A classical solution concept in fully observable nondeterministic (FOND) planning, is the strong policy (aka winning strategy in the closely related area of reactive synthesis), i.e., such a policy ensures that the goal is reached in an adversarial environment. When strong policies are not available or there is no evidence that the environment is adversarial, one can resort to best-effort policies, which always exist, and which follow the classic decision-theoretic principle that an agent should not use a dominated strategy. A typical positional best-effort policy works as follows: from every state, it follows a strong policy if one exists from that state (such states are called ``strong-winning''), else a weak policy if one exists from that state (``weak-winning''), and else is unconstrained (``losing''). In this work, we introduce a sound and complete planner for both best-effort planning and strong planning. The algorithm that underpins the planner is quite simple: it represents certain sets of states, such as the winning regions, by their $\subseteq$-minimal elements. The algorithm returns uniform policies, i.e., it returns a policy $\pi_t$ that is a strong solution starting in every strong-winning state, and it returns a policy $\pi_w$ that is a weak solution starting in every weak-winning state, and it provides a certificate for the set of losing states. We implemented the algorithm with some simple optimizations (calling it FONDANT), and evaluated it on a benchmark set consisting of the instances that were used in the evaluation of leading strong planners PR2 and FOND-SAT, and the best-effort planner BeSyftP. On coverage, our implementation is at least as good on all domains, and outperforms on some domains; and on wall time, it is slower on small and medium-sized instances, and outperforms on larger instances.