Score
Design and analyze rigorous impossibility or resource-certification proofs that show no algorithm, estimator, or protocol can surpass a stated threshold for resources (time, queries, samples, memory, rounds, error, or regret) under a given model. This work constructs adversarial instances, reductions, information-theoretic and combinatorial arguments, and specialized constructions from representation, spectral, or proof-complexity techniques (e.g., resolution or graph-spectrum arguments), and adapts those methods to settings such as deterministic vs randomized, adaptive vs non‑adaptive, differential-privacy variants (approximate‑DP, continual), streaming/finite‑time or finite‑dimensional regimes, and exponential-time or other asymptotic regimes.
This paper addresses the efficient verification of black-box model properties under a two-prover setting—one honest and one adversarial. We propose “Referee Learning,” a framework that enables reliable inference from intractably samplable distributions via game-theoretic interaction, distributional sampling, and a polynomial-communication protocol—requiring only a single query to the honest oracle. Our key contribution is the first integration of a two-prover mechanism into model selection, overcoming fundamental sample-complexity limitations of single-prover approaches. For any dimension $d$, our method achieves $(1+varepsilon)$-approximation to the optimal model with communication cost $(1+1/varepsilon^2)cdotmathrm{poly}(d)$ bits, enabling ultra-low-cost, high-accuracy model discrimination.
This paper studies sublinear-time property testing under the online adversarial erasure model, where an adversary may erase up to $t$ input positions after each query, dynamically obstructing data access. We formally define this online erasure query model and establish a systematic testing framework integrating probabilistic verification, adaptive querying, and structured violation analysis. Our main contributions are: (1) proving that linearity and quadraticity are efficiently testable with tight query complexity $Theta(log t)$, matching the non-erasure setting for constant $t$; (2) demonstrating that fundamental properties such as sortedness are inherently untestable under online erasures; and (3) providing the first characterization of the intrinsic vulnerability of multiple property classes to dynamic data corruption, thereby laying a theoretical foundation for designing robust sublinear algorithms resilient to online adversarial interference.
This work addresses the challenge of efficiently verifying the validity of statistical learning models on a given data distribution without trusting the learner. It introduces publicly verifiable certificates of statistical validity (pvCSVs), establishing the first non-interactive, publicly verifiable proof system for learning that applies to adaptive statistical query (SQ) algorithms. The proposed framework enables any user to verify a model’s performance on their own distribution with a sample complexity of only $O(\log k)$, a significant improvement over the $\tilde{O}(\sqrt{k})$ complexity of standard SQ learning algorithms. This result provides a systematic characterization of the capabilities and limitations of the SQ model in the context of verifiable learning.
This paper studies online learning against adaptive adversaries under differential privacy constraints, addressing both the realizable and agnostic settings. Building on Littlestone dimension theory, we propose the first differentially private algorithm that simultaneously ensures strong performance guarantees: it achieves an $O_d(log T)$ mistake bound in the realizable setting and a $ ilde{O}_d(sqrt{T})$ sublinear regret in the agnostic setting. Our work closes a fundamental gap in the optimization of mistake bounds for private online learning against adaptive adversaries and constitutes the first systematic extension of differentially private online learning to the agnostic setting. Crucially, we provide a rigorous proof that all concept classes with finite Littlestone dimension are learnable under differential privacy in both settings. This result substantially advances the theoretical foundations of private online learning, unifying and generalizing prior analyses limited to the realizable case or oblivious adversaries.
This work addresses the problem of recovering hidden additive structure from a set perturbed by adversarial noise, where the symmetric difference between the observed and true sets is bounded yet potentially destructive to the original structure. We introduce the first composable Balog–Szemerédi–Gowers (BSG) compiler interface that requires no prior commitments, supports persistent membership queries, admits exact finite sampling, and preserves sharp quality guarantees. Our approach integrates an algorithmic polynomial Freiman–Ruzsa theorem without size priors, deterministic amplification, and adaptive product sampling. When η = O(K⁻¹/²), the algorithm outputs a subspace V in FPT time such that |V| ≤ |A| and 𝒩_V(A) ≤ K^O(1). For any η < 1, it produces—in polynomial sample complexity and XP time—a common list covering all hidden sets compatible with the observations.
This work addresses the challenge of formally verifying or automatically refuting ε-differential privacy (ε-DP) in complex mechanisms involving both discrete and continuous random sampling. To this end, the authors propose a fully automated refutation method based on upper-expectation supermartingales and lower-expectation submartingales. The approach simultaneously searches for a pair of inputs violating ε-DP and a non-negative output function that witnesses the violation, leveraging expected divergence to construct a sound and semi-complete proof rule. This is the first technique to enable fully automated ε-DP refutation with guarantees of soundness and semi-completeness while supporting mixed distributions. The prototype tool, SuperDP, outperforms existing methods on several challenging benchmarks, successfully disproving ε-DP for mechanisms previously beyond the reach of automated analysis.
This study investigates whether scarcity alone is sufficient to defend against Sybil attacks, revealing that the structural properties of resources critically determine the cost of concentrating influence. By formally modeling adversarial cost as \(C(s,T)\), the work introduces the notion of "influence amortization" and demonstrates that when resources are divisible, exhibit additive influence, are temporally reusable, or permit identity transfer, an adversary can reduce the cost below linear—specifically, \(C(s,T) = o(sT)\). Conversely, resources that are throughput-constrained, non-transferable, and locally bounded in time achieve a lower bound of \(C(s,T) = \Omega(sT)\). This establishes a fundamental asymptotic separation between two classes of resources, providing theoretical limits and design principles for Sybil-resistant mechanisms.
This work systematically translates classical impossibility results—such as Turing undecidability, Arrow’s impossibility theorem, and the no-free-lunch theorem—into quantifiable design principles for trustworthy AI. It introduces the notion of a “deterministic horizon,” an accuracy upper bound computable a priori solely from model depth and embedding width. By integrating information theory, circuit complexity, and residual stream analysis, the study reveals fundamental capability limits in multi-stage retrieval and preference learning. Empirical evaluation across twelve Transformer architectures shows horizons between 19–31 layers, with optimal fine-tuning recovering less than 4% accuracy. Furthermore, quantified neural inference incurs zero-knowledge verification overheads of 110–190× per nonlinear activation. The paper culminates in a framework comprising sixteen paired specifications, two of which have already been formally proven in combination.