An Analytical Approach to Privacy and Performance Trade-Offs in Healthcare Data Sharing

📅 2025-08-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the trade-off between privacy preservation and model utility in healthcare data sharing, with a focus on vulnerable subpopulations—such as elderly patients, frequent hospitalizers, and racial/ethnic minorities—who face heightened re-identification risk. Using real-world inpatient data for sepsis prediction, we systematically evaluate three anonymization techniques—k-anonymity, Zheng’s method, and MO-OBAM—across demographic and healthcare utilization attributes, measuring both privacy risk and machine learning performance (accuracy/recall). Results show that MO-OBAM substantially outperforms conventional approaches: it achieves strong privacy guarantees while incurring only ~2% degradation in model performance. Crucially, we uncover a dual characteristic of vulnerable subpopulations—they exhibit both elevated privacy risk and high predictive importance for downstream models. Building on this insight, we propose a subpopulation-aware anonymization optimization framework, offering a practical, responsible trade-off paradigm for equitable and privacy-preserving healthcare data sharing.

Technology Category

Machine Learning: PrivacyData Mining & Knowledge Management: Anomaly/Outlier DetectionComputer Vision: Bias, Fairness & Privacy

Application Category

User Modeling, Personalization and Recommendation: User privacy protection in personalized systemsSecurity and Privacy: Data transparency and provenanceResponsible Web: Data and user privacy-enhancing technologies for the Web
📝 Abstract
The secondary use of healthcare data is vital for research and clinical innovation, but it raises concerns about patient privacy. This study investigates how to balance privacy preservation and data utility in healthcare data sharing, considering the perspectives of both data providers and data users. Using a dataset of adult patients hospitalized between 2013 and 2015, we predict whether sepsis was present at admission or developed during the hospital stay. We identify sub-populations, such as older adults, frequently hospitalized patients, and racial minorities, that are especially vulnerable to privacy attacks due to their unique combinations of demographic and healthcare utilization attributes. These groups are also critical for machine learning (ML) model performance. We evaluate three anonymization methods-$k$-anonymity, the technique by Zheng et al., and the MO-OBAM model-based on their ability to reduce re-identification risk while maintaining ML utility. Results show that $k$-anonymity offers limited protection. The methods of Zheng et al. and MO-OBAM provide stronger privacy safeguards, with MO-OBAM yielding the best utility outcomes: only a 2% change in precision and recall compared to the original dataset. This work provides actionable insights for healthcare organizations on how to share data responsibly. It highlights the need for anonymization methods that protect vulnerable populations without sacrificing the performance of data-driven models.
Problem

Research questions and friction points this paper is trying to address.

Balancing privacy and utility in healthcare data sharing
Identifying vulnerable populations at risk from privacy attacks
Evaluating anonymization methods for data protection and ML performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluates three anonymization methods for privacy
Identifies vulnerable sub-populations for targeted protection
MO-OBAM method maintains high ML utility performance