SynSUM - Synthetic Benchmark with Structured and Unstructured Medical Records

📅 2024-09-13
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
Joint modeling of structured and unstructured clinical data remains challenging due to scarcity of real-world data and stringent privacy constraints. Method: We propose the first synthetic data generation framework integrating causal knowledge with large language models (LLMs). Leveraging an expert-defined causal Bayesian network, we generate 10,000 structured background variables (e.g., symptoms, diagnoses) for respiratory disease patients; concurrently, GPT-4o produces high-fidelity, semantically consistent clinical notes aligned with these variables. Contribution/Results: Our dataset achieves explicit causal alignment between structured variables and unstructured text—a first in the literature. Empirical evaluation demonstrates strong performance on clinical information extraction, multimodal reasoning, and causal inference tasks. It has undergone expert quality assessment and enables reproducible research: we publicly release the first benchmark and baseline models supporting structured–unstructured joint modeling.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageMachine Learning: Large Multimodal Models (LMMs)Reasoning under Uncertainty: Causality

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsWeb Mining and Content Analysis: Bridging structured and unstructured data
📝 Abstract
We present the SynSUM benchmark, a synthetic dataset linking unstructured clinical notes to structured background variables. The dataset consists of 10,000 artificial patient records containing tabular variables (like symptoms, diagnoses and underlying conditions) and related notes describing the fictional patient encounter in the domain of respiratory diseases. The tabular portion of the data is generated through a Bayesian network, where both the causal structure between the variables and the conditional probabilities are proposed by an expert based on domain knowledge. We then prompt a large language model (GPT-4o) to generate a clinical note related to this patient encounter, describing the patient symptoms and additional context. We conduct both an expert evaluation study to assess the quality of the generated notes, as well as running some simple predictor models on both the tabular and text portions of the dataset, forming a baseline for further research. The SynSUM dataset is primarily designed to facilitate research on clinical information extraction in the presence of tabular background variables, which can be linked through domain knowledge to concepts of interest to be extracted from the text - the symptoms, in the case of SynSUM. Secondary uses include research on the automation of clinical reasoning over both tabular data and text, causal effect estimation in the presence of tabular and/or textual confounders, and multi-modal synthetic data generation.
Problem

Research questions and friction points this paper is trying to address.

Facilitates research on clinical information extraction.
Links unstructured clinical notes to structured variables.
Supports automation of clinical reasoning and causal effect estimation.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synthetic dataset links structured and unstructured medical records
Bayesian network generates tabular data with expert-defined probabilities
GPT-4 generates clinical notes from structured patient data
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Ghent University
P
P. Rabaey
IDLab, Department of Information Technology, Ghent University - imec, Ghent, Belgium
H
Henri Arno
IDLab, Department of Information Technology, Ghent University - imec, Ghent, Belgium
S
Stefan Heytens
Department of Public Health and Primary Care, Ghent University, Ghent, Belgium
Thomas Demeester
Thomas Demeester
Associate professor, Ghent University - imec
Artificial IntelligenceNatural Language Processing(past: electromagnetics)