Revealing After Overwriting: An Exponential POMDP OPE Lower Bound under History-Dependent Logging

๐Ÿ“… 2026-10-04
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the exponential sample complexity and evidence separation challenges in offline policy evaluation (OPE) for partially observable Markov decision processes (POMDPs) under history-dependent logging policies. It develops a theoretical framework that elucidates state decodability and target evidence separation mechanisms, proposing finite-class OPE guarantees based on common observation representations. By introducing reset erasure techniques, the work establishes an exponential lower bound, while a novel post-reset calibration strategy is devised to eliminate time-step dependence. Through Kullbackโ€“Leibler divergence analysis and readout rate derivations, this research confirms the exponential sample complexity lower bound and achieves polynomial-dependence guarantees on second-moment costs. Ultimately, it attains linear sample efficiency independent of the planning horizon, offering a principled resolution to the fundamental statistical bottlenecks inherent in history-dependent POMDP evaluation.
๐Ÿ“ Abstract
Multi-step revealing can make off-policy evaluation tractable under memoryless logging. With history-dependent logging, state decodability and target-relevant evidence can separate. For every horizon $H\ge3$, we construct two exactly realizable POMDPs with four actions, at most four states per layer, a known logger, and a memoryless target. Action overlap, history coverage, and observation-only revealing remain bounded independently of $H$, yet the target values differ by $1/2$ and the KL divergence between the logged laws is $\Theta(4^{-(H-1)})$, forcing exponential sample complexity. Logger memory makes states distinguishable, while reset erases the model-distinguishing evidence preserved by the target. A separate construction retains this barrier with common, known observation-only revealing operators. Under action and history coverage, we give a finite-class OPE guarantee using common observable value representations that remain valid at every history. The sample bound depends polynomially on their second-moment cost. In the common-operator construction, the same value direction has constant marginal decoding cost but exponential history-conditioned cost. Finally, on a fixed four-action continuum, we derive matching passive and budgeted readout rates. With one known channel and unit read cost, early reads are optimal. With unknown sensor bias, early reads alone remain exponentially costly. Combining them with post-reset calibration gives sample complexity independent of $H$ when both read types receive fixed positive expected budgets per trajectory.
Problem

Research questions and friction points this paper is trying to address.

Off-policy evaluation
POMDP
History-dependent logging
Sample complexity
Exponential lower bound
Innovation

Methods, ideas, or system contributions that make the work stand out.

Off-Policy Evaluation
POMDP
History-Dependent Logging
Exponential Lower Bound
Sample Complexity
๐Ÿ”Ž Similar Papers