Score
Designs and analyzes estimation and inference procedures for semiparametric models that satisfy local differential privacy constraints; this includes constructing locally-private estimators, influence-function–based de-biased estimators, confidence intervals, and tests that aim to preserve semiparametric properties such as asymptotic efficiency, unbiasedness, and double-robust or rate-double-robust convergence rates under local privacy. Practitioners develop privatization mechanisms and modified estimating equations or de-biasing steps to transfer semiparametric guarantees into the locally-private setting.
This study addresses the challenge of performing efficient and robust statistical inference for target parameters—such as causal effects—under local differential privacy constraints. We propose a novel integration of the rate-double-robust inference framework with local privacy mechanisms, where noise is injected into individual-level data to ensure privacy protection. Leveraging semiparametric theory, our approach successfully transfers desirable properties of non-private estimators to the privatized setting. The resulting method guarantees unbiasedness and achieves semiparametric efficiency, while also preserving the original estimator’s convergence rate under privacy perturbations. Furthermore, it maintains favorable asymptotic performance under both nonparametric and parametric perturbation regimes, demonstrating broad applicability and robustness in practical privacy-preserving inference tasks.
Existing differentially private statistical inference methods often neglect the impact of data clamping on the sampling distribution of estimators, leading to undercoverage of confidence intervals and uncontrolled Type-I error rates in hypothesis testing. To address this, we propose a debiased parametric bootstrap framework that—novelty—integrates indirect estimation with adaptive simulation to invert the clamping mechanism, yielding consistent and asymptotically minimum-variance estimators of the original parameters. Our approach obviates explicit modeling of clamping thresholds and applies broadly to canonical settings including location-scale normal models, linear regression, and logistic regression. We establish theoretical guarantees that the method exactly calibrates confidence interval coverage probabilities and hypothesis test significance levels under differential privacy. Empirical evaluations demonstrate substantial finite-sample improvements over state-of-the-art baselines.
Under existing differential privacy (DP) frameworks, there is a lack of general-purpose statistical inference methods—particularly when privately releasing multiple bootstrap estimates to construct confidence intervals (CIs), where privacy cost accumulation remains intractable and theoretical guarantees for sampling distribution inference are absent. This paper introduces DP Bootstrap, a novel paradigm: (i) it establishes the first universal privacy cost analysis for a single DP bootstrap release; (ii) it proposes a numerical composition method to precisely aggregate privacy budgets across multiple releases; (iii) it achieves asymptotically optimal privacy guarantees within the Gaussian DP (GDP) framework; and (iv) it pioneers DP inference for quantile regression. Theoretically, the resulting CIs attain nominal coverage, while point estimators enjoy consistency, asymptotic efficiency, and minimax-optimal convergence rates. Empirical evaluation on the 2016 Canadian Census data demonstrates significant improvements over baselines in mean estimation, logistic regression, and quantile regression.
This paper investigates semi-feature privacy under local differential privacy (LDP), where a subset of features is publicly released while the remaining features and labels must satisfy LDP constraints. We formally introduce the “semi-feature LDP” model—the first principled framework deviating from conventional full-feature perturbation. For nonparametric regression, we propose HistOfTree, an estimator that integrates histogram-based partitioning with adaptive tree-structured feature splitting, augmented by a data-driven hyperparameter selection strategy. We establish its minimax-optimal convergence rate, strictly improving upon existing LDP lower bounds for analogous problems. Extensive experiments on synthetic and real-world datasets demonstrate consistent and significant performance gains over state-of-the-art methods. Our core contributions unify conceptual modeling innovation, algorithmic design, and theoretical advancement—establishing both a new privacy paradigm and provably optimal estimation under semi-feature LDP.
This paper investigates the fundamental statistical estimation limits under user-level local differential privacy (LDP) in the multi-observation setting: each of $n$ users holds $T$ independent observations. The authors establish the first general information-theoretic lower bound for user-level LDP, revealing a $T$-driven phase transition in both mean estimation and nonparametric density estimation—namely, a critical threshold exists below which estimation risk does not vanish with $n$, and above which consistent estimation becomes possible. They further demonstrate that high-dimensional sparse mean estimation is feasible under user-level LDP but impossible under standard item-level LDP. Tight (up to logarithmic factors) minimax upper and lower bounds are derived for univariate/multivariate mean estimation, sparse mean estimation, and density estimation, explicitly characterizing the critical interplay among $T$, dimension $d$, and sparsity $s$. These results provide novel feasibility criteria for statistical inference under user-level privacy constraints.
This work presents the first unbiased estimator and self-normalized inference procedure for online quantile regression under user-level ε-local differential privacy (LDP). To address the challenge that the server cannot directly access raw estimating equations, the authors propose a mechanism based on a finite-alphabet communication channel: each user uploads a single privatized contribution by combining support-aware randomized quantization with randomized response, and the server employs a public decoder to correct bias, reconstruct an unbiased gradient, and perform online inference via projected Polyak–Ruppert averaging. The method avoids Hessian computation and establishes asymptotic normality for pre-specified scalar contrasts. Theoretical analysis guarantees privacy, unbiasedness, consistency, and asymptotic normality. Empirical results demonstrate superior privacy–utility trade-offs compared to Laplace and high-dimensional exponential mechanisms, with performance approaching the non-private benchmark as the privacy budget increases.
This work addresses the high sample complexity inherent in estimating monotone statistics under differential privacy. The authors propose an improved subsample-and-aggregate algorithm that introduces a tunable parameter \( t \) to achieve a controllable trade-off between runtime and sample complexity. While maintaining polynomial time complexity, the method reduces the required sample size by a factor of \( t \) compared to conventional approaches and is nearly optimal in terms of query complexity. Empirical evaluations demonstrate that the algorithm substantially enhances performance in private estimation tasks, including eigenvalue estimation, loss estimation, and single-parameter estimation in high-dimensional models.
This work addresses the utility degradation commonly observed in local differential privacy (LDP) when publishing functional data, such as trajectories or density curves, due to excessive perturbation. It introduces Geo-Privacy into the local model for the first time, adaptively calibrating privacy guarantees based on the distance between functions: stronger protection is provided for nearby functions, while sufficiently distant functions remain distinguishable. This approach overcomes the utility bottleneck inherent in conventional LDP mechanisms for releasing continuous functions. By integrating continuous function modeling with an efficient perturbation strategy, the proposed mechanism demonstrably enhances data utility while preserving strong privacy for neighboring functions, as validated across multiple real-world datasets.
This study addresses the challenge that noise introduced by differential privacy severely impedes effective statistical inference on large-scale privacy-preserving data, as existing approaches often rely on strong parametric assumptions or lack scalability. We propose a two-step approximate Bayesian inference framework: first imputing the differentially private data, then sampling from the non-private posterior distribution. The method achieves asymptotic validity under weak assumptions in large samples while simultaneously satisfying conservative frequentist properties, thereby combining Bayesian flexibility with frequentist reliability. Building upon and refining the approach of Guha and Reiter (2025), we demonstrate the method’s effectiveness and practical utility through simulation studies and an analysis of homeownership using the 2022 American Community Survey.