🤖 AI Summary
This study addresses the prohibitive cost of parallel sampling evaluation in masked generative models caused by unknown token dependencies. To overcome this, we propose a counterfactual probing mechanism based on an approximate conditional oracle tailored for hidden forest-structured distributions, which identifies safe batches for concurrent generation. By integrating Hellinger error bound analysis with tunable parameter trade-off strategies, our approach enables efficient parallel sampling without requiring full recovery of the dependency tree. Theoretically, we prove that both the total number of evaluations and the sequential depth can surpass the linear constraint imposed by sequence length, achieving sublinear complexity. Furthermore, we establish rigorous theoretical lower bounds to guarantee these results. This work provides a solid theoretical foundation for accelerating inference in masked generative models.
📝 Abstract
Masked generative models offer parallel token prediction, but accurate parallel sampling must account for dependencies among tokens. When dependencies are unknown, finding safe batches also costs model evaluations. We study whether total evaluations, including discovery, can be sublinear in sequence length $N$; sublinear sequential depth then follows. We consider discrete distributions with hidden forest structure, accessed through a fixed approximate conditional oracle. Under explicit regularity conditions and uniform Hellinger error bounds, for any fixed target accuracy $\varepsilon\in (0,1/8]$ and sufficiently large $N$, our sampler achieves seed-averaged total-variation error at most $\varepsilon$, with total masked-state submissions and sequential depth both bounded by $O(N^C \varepsilon^a)$ for constants $0 <C <1$ and $a > 0$. These guarantees use polynomial vocabulary size and an edge-response lower bound set by $N$ and $\varepsilon$. The sampler shares evaluations of hypothetical reveals across dependence tests to identify safe parallel batches without requiring full recovery of the hidden forest. A tunable parameter trades probing cost against irreversible commit rounds. In the same class, any admissible irreversible product-commit sampler attaining the same seed-averaged accuracy requires $Ω(N^c \varepsilon^b)$ counterfactual submissions or commit rounds in the worst case, for constants $c,b>0$.