Estimation of discrete distributions in relative entropy, and the deviations of the missing mass

📅 2025-04-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses high-probability estimation of discrete distributions over a finite alphabet under the Kullback–Leibler (KL) divergence, focusing on the sparse regime where the sample size is smaller than the alphabet size. We propose a data-driven, adaptive smoothing estimator tailored to this setting. Our main contributions are threefold: First, we establish the first non-asymptotic, high-probability upper bound on the KL risk, which is minimax optimal up to constant factors. Second, we derive a sharp high-probability upper bound on the missing mass. Third, we prove that the KL risk of the Laplace estimator admits tight high-probability upper and lower bounds; that the minimax high-probability risk incurs an extra logarithmic factor; and—under a sparsity assumption—that the risk bound depends only on two effective sparsity parameters, enabling automatic adaptation to underlying distribution structure. Collectively, these results unify and substantially extend the theoretical foundations of classical smoothing methods, delivering rigorous, practical risk guarantees for small-sample discrete distribution estimation.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationReasoning under Uncertainty: Stochastic OptimizationSearch and Optimization: Mixed Discrete/Continuous Search

Application Category

Graph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsSecurity and Privacy: Large-scale security measurementsUser Modeling, Personalization and Recommendation: User privacy protection in personalized systems
📝 Abstract
We study the problem of estimating a distribution over a finite alphabet from an i.i.d. sample, with accuracy measured in relative entropy (Kullback-Leibler divergence). While optimal expected risk bounds are known, high-probability guarantees remain less well-understood. First, we analyze the classical Laplace (add-$1$) estimator, obtaining matching upper and lower bounds on its performance and showing its optimality among confidence-independent estimators. We then characterize the minimax-optimal high-probability risk achievable by any estimator, which is attained via a simple confidence-dependent smoothing technique. Interestingly, the optimal non-asymptotic risk contains an additional logarithmic factor over the ideal asymptotic risk. Next, motivated by scenarios where the alphabet exceeds the sample size, we investigate methods that adapt to the sparsity of the distribution at hand. We introduce an estimator using data-dependent smoothing, for which we establish a high-probability risk bound depending on two effective sparsity parameters. As part of the analysis, we also derive a sharp high-probability upper bound on the missing mass.
Problem

Research questions and friction points this paper is trying to address.

Estimating discrete distributions with relative entropy accuracy
Analyzing Laplace estimator's optimality and performance bounds
Developing adaptive estimators for sparse large-alphabet distributions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Laplace estimator with confidence-independent optimality
Minimax-optimal risk via confidence-dependent smoothing
Data-dependent smoothing for sparse distributions
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jaouad Mourtada