Efficient Learning of Truncated Boolean Product Distributions: Influence to the Rescue

📅 2026-07-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the problem of efficiently learning the natural parameters of a Boolean product distribution truncated to a subset $S$, overcoming limitations of prior approaches that require strong connectivity or anti-concentration assumptions on $S$ and suffer from exponential sample complexity. By analyzing the geometric structure of the truncation set $S$ under the distribution $\mu_z$, the paper introduces Boolean function influence theory into truncated distribution learning for the first time, generalizing the fatness condition and enabling efficient parameter estimation without relying on any parametric sampling assumptions. The proposed method achieves a sample complexity of $O(\log n / \varepsilon^2)$ in $\ell_\infty$ norm, matching the minimax rate in the untruncated setting, and establishes lower bounds that reveal an exponential dependence on both the model width and the minimal distance of the set $S$.
📝 Abstract
Learning the natural parameters $z \in \mathbb{R}^n$ of discrete distributions $μ_z$ from independent samples constrained to a subset $S \subseteq \{0,1\}^n$ is a foundational challenge in high-dimensional statistics. Existing methods for efficiently estimating truncated Boolean product distributions, notably the work of [Fotakis et al' COLT'20, Algorithmica '22], require either strong local connectivity assumptions on $S$ -- a property denoted fatness -- or stringent anti-concentration assumptions and necessitate the total mass of the truncation set to be a constant with respect to $n$. Moreover, the results in [Fotakis et al' COLT'20, Algorithmica '22] suffer from sample complexities that scale as $Ω(2^n)$ if the mass of $S$ is exponentially small in $n$. In this work, we circumvent these limitations by analyzing the geometry of $S$ under the measure $μ_z$. We refine the existing parameter estimation guarantees under the fatness assumption, improving the prior sample complexity to $O( \log n / ε^2)$ for $\ell_\infty$-recovery, matching the untruncated minimax rate. We further generalize fatness using the notion of influence utilized in the analysis of Boolean functions and provide sufficient conditions for efficient inference. Notably, unlike previous work, our method does not require sampling at arbitrary parameterizations of the model. Lastly, we establish a theoretical lower bound demonstrating the sample complexity exhibits an intrinsic exponential dependence on the width of the model and the minimum distance between elements in the set.
Problem

Research questions and friction points this paper is trying to address.

truncated Boolean product distributions
parameter estimation
sample complexity
influence
high-dimensional statistics
Innovation

Methods, ideas, or system contributions that make the work stand out.

truncated Boolean product distributions
influence
fatness
sample complexity
parameter estimation
🔎 Similar Papers
No similar papers found.