🤖 AI Summary
Addressing the challenge of balancing privacy preservation and machine learning (ML) utility in structured data sharing, this paper proposes a multi-objective optimization-driven anonymization framework. The method unifies categorical variable modeling, multi-dimensional utility assessment, and cross-dataset robustness validation within a single optimization paradigm—first of its kind. It integrates the NSGA-II evolutionary algorithm, an extended hybrid privacy metric combining k-anonymity and l-diversity, an ML-performance-sensitive information loss function, and quantitative risk evaluation against linkage and homogeneity attacks. Evaluated on multiple benchmark datasets, the framework achieves an average 12.7% reduction in information loss, up to a 38.4% decrease in the number of attack-vulnerable individuals, and ML model accuracy within 1.5 percentage points of that attained on original data—outperforming state-of-the-art anonymization approaches.
📝 Abstract
Data is essential for secondary use, but ensuring its privacy while allowing such use is a critical challenge. Various techniques have been proposed to address privacy concerns in data sharing and publishing. However, these methods often degrade data utility, impacting the performance of machine learning (ML) models. Our research identifies key limitations in existing optimization models for privacy preservation, particularly in handling categorical variables, assessing data utility, and evaluating effectiveness across diverse datasets. We propose a novel multi-objective optimization model that simultaneously minimizes information loss and maximizes protection against attacks. This model is empirically validated using diverse datasets and compared with two existing algorithms. We assess information loss, the number of individuals subject to linkage or homogeneity attacks, and ML performance after anonymization. The results indicate that our model achieves lower information loss and more effectively mitigates the risk of attacks, reducing the number of individuals susceptible to these attacks compared to alternative algorithms in some cases. Additionally, our model maintains comparative ML performance relative to the original data or data anonymized by other methods. Our findings highlight significant improvements in privacy protection and ML model performance, offering a comprehensive framework for balancing privacy and utility in data sharing.