On the Optimal Number of Grids for Differentially Private Non-Interactive $K$-Means Clustering

📅 2026-03-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical limitation in existing non-interactive differentially private $k$-means clustering methods, which lack theoretical guidance in grid discretization and consequently struggle to balance quantization bias against noise-induced perturbation. The paper introduces, for the first time, a theoretically grounded criterion for selecting grid resolution by minimizing an upper bound on the expected distortion of the $k$-means objective function. This criterion explicitly links the number of clusters, dataset size, and privacy budget. Building upon histogram-based private data summaries, the proposed method integrates rigorous quantization error analysis with optimization to significantly enhance clustering accuracy under strict differential privacy constraints, outperforming state-of-the-art non-interactive approaches.

Technology Category

Machine Learning: PrivacyConstraint Satisfaction and Optimization: Distributed CSP/OptimizationSearch and Optimization: Distributed Search

Application Category

Security and Privacy: Data transparency and provenanceUser Modeling, Personalization and Recommendation: User privacy protection in personalized systemsEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasets
📝 Abstract
Differentially private $K$-means clustering enables releasing cluster centers derived from a dataset while protecting the privacy of the individuals. Non-interactive clustering techniques based on privatized histograms are attractive because the released data synopsis can be reused for other downstream tasks without additional privacy loss. The choice of the number of grids for discretizing the data points is crucial, as it directly controls the quantization bias and the amount of noise injected to preserve privacy. The widely adopted strategy selects a grid size that is independent of the number of clusters and also relies on empirical tuning. In this work, we revisit this choice and propose a refined grid-size selection rule derived by minimizing an upper bound on the expected deviation in the K-means objective function, leading to a more principled discretization strategy for non-interactive private clustering. Compared to prior work, our grid resolution differs both in its dependence on the number of clusters and in the scaling with dataset size and privacy budget. Extensive numerical results elucidate that the proposed strategy results in accurate clustering compared to the state-of-the-art techniques, even under tight privacy budgets.
Problem

Research questions and friction points this paper is trying to address.

differential privacy
K-means clustering
non-interactive
grid discretization
privacy-preserving
Innovation

Methods, ideas, or system contributions that make the work stand out.

differentially private clustering
non-interactive K-means
grid discretization
privacy-utility tradeoff
histogram-based synopsis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Gokularam Muthukrishnan
Center of Data for Public Good, Foundation for Science Innovation and Development, Indian Institute of Science, Bengaluru 560012, India
Anshoo Tandon
Anshoo Tandon
CDPG (FSID, IISc)
Information TheoryCoding TheorySignal ProcessingPrivacy