SKALD: Scalable K-Anonymisation for Large Datasets

📅 2025-05-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing k-anonymization tools (e.g., ARX) process large-scale tabular data in isolated blocks under memory constraints, failing to guarantee global k-anonymity and often degrading utility or even failing outright. Method: We propose the first scalable k-anonymization framework based on block-wise statistical aggregation. It employs incremental statistic extraction, inter-block coordination of generalization parameters, and memory-aware streaming processing to enforce *global* k-anonymity—rather than merely intra-block anonymity—on terabyte-scale datasets using limited RAM. Contribution/Results: Experiments demonstrate that our approach achieves up to several-fold speedup over ARX and other baselines, while improving post-anonymization classification accuracy by an average of 12.7%. It simultaneously ensures rigorous privacy guarantees, high data utility, and linear scalability—addressing the long-standing tension among privacy, utility, and efficiency in large-scale anonymization.

Technology Category

Data Mining & Knowledge Management: Scalability, Parallel & Distributed SystemsMachine Learning: PrivacySearch and Optimization: Distributed Search

Application Category

Security and Privacy: Large-scale security measurementsResponsible Web: Ethical and legal aspects of web-scale data analysis, uses and collection practicesGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphs
📝 Abstract
Data privacy and anonymisation are critical concerns in today's data-driven society, particularly when handling personal and sensitive user data. Regulatory frameworks worldwide recommend privacy-preserving protocols such as k-anonymisation to de-identify releases of tabular data. Available hardware resources provide an upper bound on the maximum size of dataset that can be processed at a time. Large datasets with sizes exceeding this upper bound must be broken up into smaller data chunks for processing. In these cases, standard k-anonymisation tools such as ARX can only operate on a per-chunk basis. This paper proposes SKALD, a novel algorithm for performing k-anonymisation on large datasets with limited RAM. Our SKALD algorithm offers multi-fold performance improvement over standard k-anonymisation methods by extracting and combining sufficient statistics from each chunk during processing to ensure successful k-anonymisation while providing better utility.
Problem

Research questions and friction points this paper is trying to address.

Addressing k-anonymisation for large datasets with limited RAM
Improving performance over standard per-chunk k-anonymisation methods
Ensuring data privacy while maintaining better utility
Innovation

Methods, ideas, or system contributions that make the work stand out.

SKALD algorithm for large dataset k-anonymisation
Extracts sufficient statistics from data chunks
Ensures k-anonymisation with limited RAM
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.