On Computational Hardness of Mistake-Bounded Language Generation: A Random-Oracle Query Separation

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether language classes that are information-theoretically learnable without error remain computationally hard under a bounded number of queries. Working within the random oracle model, the authors construct a countably infinite family of languages with closure dimension zero and, by integrating a mistake-bounded generation framework with probabilistic methods, demonstrate that any polynomial-query generator incurs an exponentially growing worst-case expected number of errors for sufficiently large input lengths. This work establishes the first separation in the random oracle model between information-theoretic tractability and computational hardness under query bounds, thereby revealing a fundamental limitation imposed by computational resources on the achievable error rate in language generation tasks.
📝 Abstract
Generation in the limit guarantees eventual generation for every countable collection of infinite languages in the model of Kleinberg and Mullainathan [KM24], while closure dimension characterizes stronger information-theoretic guarantees [RLT25]. Neither restricts per-output computation. The cumulative-mistake objective in mistake-bounded generation makes finite failure prefixes quantitative [KPR26], and a per-output query budget exposes their computational source. Polynomial-time algorithms are known for parities, conjunctions, and monotone functions with polynomially many maxterms [JKO26]. We ask whether information-theoretic ease can coexist with bounded-access computational hardness. Relative to a random oracle $H$, we answer yes by constructing a countable collection $C^\star$ of infinite languages with closure dimension zero. Almost surely on the same $H$, an unbounded generator makes zero mistakes on every target and every complete distinct enumeration. Yet, writing $λ$ for the target-seed length, every fixed uniform generator $G$ with polynomially many oracle queries in $λ$ and the output index $i$ has a constant $c_G>0$ such that, for every sufficiently large $λ$, some target incurs more than $2^{c_Gλ}$ expected mistakes within its first $2(\lceil 2^{c_Gλ}\rceil+1)$ canonical outputs. Infinite accidental agreement enables exhaustive search; sparse queries hide fresh target values. Thus, in the random-oracle model, zero-mistake information-theoretic generation coexists with a generator-dependent exponential lower bound on worst-case expected mistakes under polynomial-query access.
Problem

Research questions and friction points this paper is trying to address.

mistake-bounded generation
computational hardness
information-theoretic guarantees
random oracle
language generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

mistake-bounded generation
random oracle
closure dimension
computational hardness
polynomial-query generator