Expectation in Stochastic Games with Prefix-independent Objectives

📅 2024-05-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper investigates expected-value computation in prefix-independent stochastic games. Methodologically, it introduces the first general reduction that linearly transforms expected-value optimization into the decision problem of almost-sure satisfaction of threshold Boolean objectives. It proves that optimal strategies require no more memory than the known memory upper bounds for the corresponding Boolean objectives. The paper establishes, for the first time, UP ∩ coUP complexity bounds for the expected-value decision problem under window mean-payoff objectives—both fixed-window and bounded-window variants—and provides succinct certificate verification mechanisms. The proposed unified framework is extensible to a broad class of prefix-independent quantitative objectives, balancing theoretical rigor with algorithmic tractability.

Technology Category

Reasoning under Uncertainty: Stochastic OptimizationConstraint Satisfaction and Optimization: Satisfiability Modulo TheoriesGame Theory and Economic Paradigms: Cooperative Game Theory

Application Category

Economics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystemsGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsWeb Mining and Content Analysis: Models for Web evolution
📝 Abstract
Stochastic two-player games model systems with an environment that is both adversarial and stochastic. In this paper, we study the expected value of quantitative prefix-independent objectives in stochastic games. We show a generic reduction from the expectation problem to linearly many instances of almost-sure satisfaction of threshold Boolean objectives. The result follows from partitioning the vertices of the game into so-called value classes where each class consists of vertices of the same value. Our procedure further entails that the memory required by both players to play optimally for the expectation problem is no more than the memory required by the players to play optimally for the almost-sure satisfaction problem for a corresponding threshold Boolean objective. We show the applicability of the framework to compute the expected window mean-payoff measure in stochastic games. The window mean-payoff measure strengthens the classical mean-payoff measure by computing the mean-payoff over a window of bounded length that slides along an infinite path. Two variants have been considered: in one variant, the maximum window length is fixed and given, while in the other, it is not fixed but is required to be bounded. For both variants, we show that the decision problem to check if the expected value is at least a given threshold is in UP $cap$ coUP. The result follows from guessing the expected values of the vertices, partitioning them into value classes, and proving that a unique short certificate for the expected values exists. It also follows that the memory required by the players to play optimally is no more than that in non-stochastic two-player games with the corresponding window objectives.
Problem

Research questions and friction points this paper is trying to address.

Study expected value in stochastic games with prefix-independent objectives
Reduce expectation problem to almost-sure satisfaction instances
Compute expected window mean-payoff measure in stochastic games
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reduction to almost-sure satisfaction problem
Partitioning vertices into value classes
Memory-efficient optimal strategy computation
🔎 Similar Papers
2023-09-30International Symposium on Games, Automata, Logics and Formal VerificationCitations: 0
💼 Related Jobs
No related jobs found.
CNRS & LMF | ENS Paris-Saclay | Tata Institute of Fundamental Research
L
L. Doyen
CNRS & LMF, ENS Paris-Saclay, Gif-sur-Yvette, France
P
Pranshu Gaba
Tata Institute of Fundamental Research, Mumbai, India
Shibashis Guha
Shibashis Guha
Reader, Tata Institute of Fundamental Research Mumbai
Formal methods and learningcontroller synthesisMarkov decision processesautomata theory