$(ell,δ)$-Diversity: Linkage-Robustness via a Composition Theorem

📅 2025-06-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Under multi-dataset linkage attacks, conventional ℓ-diversity fails to preserve anonymity due to its deterministic requirement on sensitive attribute diversity within equivalence classes. Method: We propose (ℓ,δ)-diversity—a probabilistic and composable robust anonymity notion that integrates differential privacy principles with equivalence-class structure. It relaxes ℓ-diversity by permitting at most a δ-fraction of equivalence classes to violate the ℓ-diversity constraint, via probabilistic modeling of sensitive attribute distributions. Contribution/Results: We prove that (ℓ,δ)-diversity satisfies composition: linking k independently anonymized datasets preserves (ℓ,δ′)-diversity, where δ′ grows only slowly with k—marking a substantial improvement over the fragility of standard ℓ-diversity under composition. Additionally, we derive a utility lower bound, enabling a more favorable trade-off between privacy guarantee strength and data utility.

Technology Category

Machine Learning: PrivacyNatural Language Processing: Safety and RobustnessData Mining & Knowledge Management: Anomaly/Outlier Detection

Application Category

Security and Privacy: Data transparency and provenanceEconomics, Online Markets and Human Computation: Fairness, privacy, and diversity in economic environmentsUser Modeling, Personalization and Recommendation: User privacy protection in personalized systems
📝 Abstract
In this paper, we consider the problem of degradation of anonymity upon linkages of anonymized datasets. We work in the setting where an adversary links together $tgeq 2$ anonymized datasets in which a user of interest participates, based on the user's known quasi-identifiers, which motivates the use of $ell$-diversity as the notion of dataset anonymity. We first argue that in the worst case, such linkage attacks can reveal the exact sensitive attribute of the user, even when each dataset respects $ell$-diversity, for moderately large values of $ell$. This issue motivates our definition of (approximate) $(ell,δ)$-diversity -- a parallel of (approximate) $(ε,δ)$-differential privacy (DP) -- which simply requires that a dataset respect $ell$-diversity, with high probability. We then present a mechanism for achieving $(ell,δ)$-diversity, in the setting of independent and identically distributed samples. Next, we establish bounds on the degradation of $(ell,δ)$-diversity, via a simple ``composition theorem,'' similar in spirit to those in the DP literature, thereby showing that approximate diversity, unlike standard diversity, is roughly preserved upon linkage. Finally, we describe simple algorithms for maximizing utility, measured in terms of the number of anonymized ``equivalence classes,'' and derive explicit lower bounds on the utility, for special sample distributions.
Problem

Research questions and friction points this paper is trying to address.

Addresses anonymity degradation in linked anonymized datasets
Proposes (ℓ,δ)-diversity to prevent sensitive attribute disclosure
Develops mechanisms to preserve diversity upon dataset linkages
Innovation

Methods, ideas, or system contributions that make the work stand out.

Introduces $(ell,delta)$-diversity for linkage robustness
Mechanism for achieving $(ell,delta)$-diversity in i.i.d. samples
Composition theorem bounds diversity degradation upon linkage