data de-identification

Applying transformations and governance (e.g., removal, pseudonymization, aggregation) to multimodal clinical data so datasets can be ethically collected, GDPR-compliant shared, and converted into structured inputs for downstream models while preserving analytic utility.

datade-identification

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

A systematic review of challenges and proposed solutions in modeling multimodal data

May 11, 2025
MF
Maryam Farhadizadeh
🏛️ University of Freiburg | University Medical Center Hamburg-Eppendorf | Ulm University

Clinical multimodal modeling faces persistent challenges including missing modalities, limited sample sizes, dimensional imbalance, and insufficient interpretability. To address these, we conduct the first structured review of 69 medical multimodal studies, establishing a “problem–solution” mapping framework, and propose guidelines for fusion strategy selection and an interpretability evaluation pathway. Methodologically, we integrate transfer learning, generative models, cross-modal attention mechanisms, and neural architecture search—emphasizing modality alignment and adaptive fusion. We distill five major challenge categories and their empirically validated solutions, yielding a comprehensive technical roadmap spanning medical imaging, genomics, wearable sensors, and electronic health records. This work provides both theoretical foundations and practical paradigms for designing, evaluating, and clinically deploying multimodal AI systems in healthcare.

Advancing methods for interpretable medical multimodal modelingIdentifying challenges in modeling multimodal clinical dataReviewing solutions for missing data and fusion techniques

Must-Read Papers

Most classic and influential ideas
View more

A Privacy-Preserving Ecosystem for Developing Machine Learning Algorithms Using Patient Data: Insights from the TUM.ai Makeathon

Oct 29, 2025
SS
Simon Süwer
🏛️ University of Hamburg | Ludwig-Maximilians-Universität München | MI4People – Non-Profit for AI for Social Good

Addressing the dual challenges of GDPR compliance and patient privacy preservation in rare-disease settings with limited clinical data. Method: We propose a privacy-by-design clinical AI modeling framework that integrates a synthetic clinical knowledge graph (cKG) for structure-preserving initial modeling and leverages the FeatureCloud federated learning platform to enable secure, on-site model training and evaluation within hospital-controlled environments. A multi-stage security protocol ensures raw data never leaves the premises—comprising cKG-based pre-modeling with content-level de-identification, end-to-end data isolation within sandboxed execution environments, automated security pipelines, and aggregation-only metric evaluation. The framework supports secure, collaborative analysis of multi-omics and heterogeneous clinical data. Contribution/Results: At TUM.ai Makeathon 2024, 50 participants successfully developed patient classification and diagnostic models without accessing any real patient data, demonstrating the framework’s feasibility, regulatory compliance, and efficiency in privacy-constrained collaborative AI development.

Developing privacy-preserving machine learning algorithms using patient dataEnabling secure AI training without exposing sensitive clinical informationOvercoming GDPR restrictions for small cohorts with rare diseases

Privacy-Aware, Public-Aligned: Embedding Risk Detection and Public Values into Scalable Clinical Text De-Identification for Trusted Research Environments

Jun 01, 2025
AC
Arlene Casey
🏛️ University of Edinburgh | University of Aberdeen | NHS Glasgow & Greater Clyde | University of Dundee

Clinical free-text reuse in trusted research environments faces dynamic privacy risk accumulation, heterogeneous identifier types, and model performance decay over time. Method: We propose a context-aware privacy risk modeling and public-value-driven hybrid de-identification framework. Integrating empirical analysis of multi-source NHS data with public value consensus, we establish a risk-stratified assessment paradigm grounded in document type, clinical context, and data flow. Our approach combines rule-based engines, context-sensitive named entity recognition (NER), temporal performance monitoring, and participatory design to yield an interpretable, traceable, and adaptive de-identification decision-support prototype. Results: Validation reveals cross-institutional and multi-diagnosis privacy risk distribution patterns, and demonstrates that evolving clinical documentation practices significantly impair model robustness. This work delivers the first empirically grounded, scalable pathway for NHS clinical text governance—balancing technical precision with auditability and regulatory compliance.

Aligning public values with scalable privacy-risk management solutionsAssessing de-identification tool performance across diverse real-world datasetsDetecting privacy risks in clinical text for secure research use

This work addresses the challenge of securely sharing electronic health records (EHRs) across institutions, which is hindered by privacy concerns, governance constraints, and interoperability limitations that impede multicenter research and medical AI development. The authors propose a novel EHR transformation framework based on irreversible geometric operators, designed under a rigorous threat model through collaborative strategy formulation between human experts and an AI agent (SciencePal). The approach integrates hybrid mechanisms tailored for high-risk scenarios and is theoretically grounded, with robustness validated against diverse privacy attacks—including reconstruction, linkage, and membership inference. Experimental results demonstrate that the transformed data effectively resist such attacks while preserving clinical interpretability and utility for machine learning tasks, thereby establishing a secure and efficient foundation for large-scale medical AI training.

clinical data interoperabilitydata sharingelectronic health records

From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine

Feb 13, 2025
LB
Lukas Buess
🏛️ Friedrich-Alexander-Universität Erlangen-Nürnberg | Technical University of Munich

Despite rapid advances, the clinical translation of generative AI in medicine remains hindered by persistent bottlenecks in transitioning from unimodal large language models to robust, clinically viable multimodal AI systems integrating medical imaging, textual narratives, and structured electronic health record data. Method: Guided by the PRISMA-ScR framework, this study conducts a systematic scoping review of 144 peer-reviewed studies published through December 2024, sourced from PubMed, IEEE Xplore, and Web of Science, with rigorous inclusion criteria. It synthesizes state-of-the-art techniques—including diffusion modeling, cross-modal alignment, explainability methods (e.g., attention visualization, saliency mapping), and clinical validation protocols. Contribution/Results: The review identifies four high-impact application domains—AI-assisted diagnosis, automated radiology/pathology reporting, generative drug discovery, and conversational clinical assistants—and delineates core challenges: data heterogeneity, limited model interpretability, ethical and regulatory uncertainties, and insufficient real-world deployment evidence. It proposes a comprehensive, trust-centered framework for responsible development and clinical integration of multimodal generative AI in healthcare.

Addressing challenges in AI integrationExploring multimodal AI in medicineReviewing AI applications in clinical settings

Despite rapid advances in multimodal large language models (MLLMs), their clinical deployment remains hindered by domain-specific bottlenecks—including scarce annotated medical data, modality bias, and limited interpretability. Method: This work systematically reviews the evolution from large language models (LLMs) to MLLMs and empirically analyzes their integration of text, medical imaging, and audio modalities for clinical decision support, radiology/pathology interpretation, patient interaction, and biomedical research. Contribution/Results: We identify three critical research directions: (1) construction of medical-domain multimodal datasets, (2) novel modality alignment techniques, and (3) an ethics-aware governance framework. Empirical evaluation demonstrates that MLLMs improve diagnostic assistance accuracy and accelerate structured reporting generation; however, performance is constrained by data scarcity, cross-modal misalignment, and opaque reasoning. This study provides both theoretical foundations and actionable pathways toward trustworthy, clinically viable MLLMs.

Addresses implementation challenges including data limitations and ethical concernsExamines MLLM applications in clinical decision support and medical imagingExplores the evolution of text-based LLMs to multimodal systems in healthcare

Latest Papers

What's happening recently
View more

This study addresses the limitations of traditional health definitions—namely their inability to capture health’s dynamic, complex, and context-dependent nature. Methodologically, it integrates physiological sensing, behavioral tracking, and environmental contextual data to construct a continuous, multidimensional framework for health representation via multimodal digital biomarkers (MDBs). The contribution is twofold: first, MDBs are theorized not merely as technical extensions of digital biomarkers but as catalysts for an ontological and epistemological shift—from health as a static state to a variable process, and from an individual attribute to a relational phenomenon. Second, the study identifies critical ethical challenges arising from this shift, including the reconfiguration of epistemic authority in data-driven healthcare, ambiguity in accountability, and gaps in governance. Accordingly, it proposes an interdisciplinary theoretical framework that balances technical feasibility with ethical robustness, offering actionable guidance for policy development, clinical implementation, and algorithmic design. (149 words)

MDBs create ontological shift by datafying health conceptsMDBs expand digital biomarkers through variability, complexity and abstractionMDBs raise ethical implications for knowledge, responsibility and governance

This study addresses the scarcity of high-quality, clinically annotated data and the constraints imposed by privacy regulations, which hinder the advancement of medical machine learning. To overcome these challenges, the authors propose a conditional generation approach leveraging large language models—specifically DeepSeek-R1, OpenBioLLM-Llama3, and Qwen 3.5—to synthesize mental health assessment reports aligned with ICD-10 coding standards. They further introduce the first multidimensional evaluation framework tailored for clinical data augmentation, which jointly assesses semantic fidelity, lexical diversity, and privacy preservation in generated texts. Experimental results demonstrate that the synthesized reports maintain clinical plausibility while effectively mitigating privacy risks, thereby substantially expanding the pool of training data available for clinical natural language processing tasks.

clinical data augmentationdata scarcitymental health

This study addresses the challenge of unreliable AI models and diminished clinical trust stemming from opaque data quality reporting in the secondary use of electronic health records (EHRs). To this end, the authors propose the first comprehensive framework for transparent data quality reporting across the entire EHR lifecycle. The framework innovatively distinguishes between data producers and consumers, explicitly defines five critical phases, and maps established data quality dimensions to specific workflow stages. Through iterative stakeholder and process analysis, a structured reporting mechanism is developed and validated on real-world datasets, demonstrating its ability to effectively trace the origins of data quality issues. The approach significantly enhances data interpretability, fitness-for-use, and governance efficacy, thereby providing a robust foundation for trustworthy AI development and clinical research.

clinical AIdata lifecycledata quality

This study addresses the challenge of balancing privacy preservation and model utility in cross-institutional sharing of radiology images and reports. The authors propose a novel de-identification pipeline that integrates a blacklist of privacy-sensitive terms, a whitelist of pathology-relevant terms, generative image filtering, and report ID removal to synthesize data that retains critical diagnostic information while eliminating personally identifiable elements. Systematic evaluation on a public chest X-ray dataset demonstrates, for the first time, that large vision–language models trained on this de-identified data achieve diagnostic performance comparable to those trained on original data, with substantially reduced re-identification risk. Furthermore, in cross-hospital transfer scenarios, combining local institutional data with the de-identified data further enhances model performance, effectively reconciling clinical utility with robust privacy safeguards.

cross-hospital transferdata utilityde-identification

This study addresses the challenge of effectively integrating patient-generated multimodal data—such as from wearable devices and self-report questionnaires—with clinical records in psychiatric practice. To bridge this gap, the authors collaborated with clinicians to design a narrative-driven dashboard that, for the first time, leverages large language models to synthesize heterogeneous data sources into coherent, context-aware natural language narratives grounded in clinical semantics, complemented by interactive visualizations. This approach substantially enhances the interpretability and clinical utility of complex data. In a user study involving 16 psychiatrists, the system demonstrated statistically significant improvements over baseline methods in revealing clinically relevant insights (p<.001) and supporting clinical decision-making (p=.004).

clinical decision-makingdata visualizationmental health

Hot Scholars

GZ

Guoying Zhao

Academy Professor, IEEE Fellow, Professor of Computer Science and Engineering, University of Oulu
Affective ComputingArtificial IntelligenceComputer VisionPattern Recognition
PL

Pierre Lison

Chief Research Scientist, Norsk Regnesentral
Natural Language ProcessingMachine LearningSpoken Dialogue SystemsMultilingual NLP
HW

Hui Wei

University of Oulu
AI SafetyTrustworthy IntelligenceComputer Vision
HY

Hao Yu

Student at CMVS, University of Oulu
computer visioncross-domain recognition