🤖 AI Summary
This work addresses the challenge of degraded data utility in local differential privacy (LDP) when collecting numerical data without prior knowledge of the data domain, which often leads to excessive clipping or over-noising. The authors propose an adaptive LDP framework in which users simultaneously upload a perturbed value and a private binary signal indicating whether their true value was clipped. By aggregating these signals across users, the system iteratively refines the clipping interval and dynamically estimates the true data domain. This approach is the first to achieve online convergence of the data domain without requiring any prior domain knowledge. Extensive experiments demonstrate that the method significantly improves accuracy across diverse datasets and LDP mechanisms, while exhibiting strong robustness to hyperparameter choices.
📝 Abstract
Local Differential Privacy (LDP) provides strong privacy guarantees for collecting numerical data. A fundamental challenge, however, is that existing LDP mechanisms require a predefined data domain, which is often unknown in practice. This lack of prior knowledge creates a critical dilemma for the data collector: if the chosen domain is too narrow, values outside the range are clipped, leading to information loss. Conversely, if the domain is too wide, excessive noise is added during the privatization process, which degrades the quality of collected data. This highlights the need for methods that can dynamically estimate the data domain.
In this work, we propose an adaptive LDP framework that addresses this problem. In our method, each user sends two pieces of information: their perturbed numerical data, and a privatized signal indicating if their original value was clipped by the current domain. By aggregating these signals, our proposed method, Adaptive Bounding of Clipping regions (ABC) method, iteratively adjusts the domain to fit the underlying data distribution without prior knowledge. Our theoretical analysis shows that the estimated data domain converges to an appropriate range.
In the empirical evaluation, the results demonstrate that our framework significantly improves the quality of numerical data collection across various datasets and underlying LDP mechanisms. We also show that the estimated range successfully converges in practice and our approach is robust to its hyperparameters through comprehensive ablation studies.