Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
While existing synthetic clinical benchmarks offer practical utility, they often suffer from structural distortions that hinder their fidelity to real-world healthcare settings. This work is the first to explicitly distinguish between—and jointly optimize—the realism and utility of synthetic data, treating utility as a constraint rather than sufficient evidence of realism. Building upon Synthea-generated data and integrating electronic health record workflows with downstream processing pipelines, the authors propose a deterministic refinement strategy that enhances realism along four dimensions: missing structure recovery, conciseness, logical consistency, and population alignment. Two iterative rounds of refinement substantially mitigate issues of data sparsity, overly concentrated distributions, and limited actionability, while consistently maintaining utility above a predefined threshold.
📝 Abstract
Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy-sensitive healthcare settings where operational data are hard to access. We study how to improve such benchmarks without breaking the downstream utility checks already used in practice. We formulate benchmark revision as utility-constrained realism improvement: dataset changes should increase realism while staying above an operational utility floor. We instantiate this idea on a care-gap benchmark derived from Synthea-generated patients exercised through demonstration electronic health record workflows and then processed by the same downstream pipeline as operational data. Realism is measured through missingness structure, simplicity, structural plausibility, and population alignment. The baseline benchmark is extremely thin: sampled-pair missingness is 79.44%, only 12.75% of rows are actionable, 38.94% of patients have zero actionable measures, and top-three token concentration reaches 100.0%. Two deterministic revisions improve these panels while remaining above the current utility floor, whereas a naive densification control preserves unrealistic templating. We further show that internal benchmark realism and source fidelity to an aggregate operational reference are related but distinct objectives. These results suggest that synthetic benchmark quality should be optimized explicitly, with utility treated as one constraint rather than as sufficient evidence of realism.
Problem

Research questions and friction points this paper is trying to address.

synthetic clinical benchmarks
realism
utility constraints
healthcare AI
data fidelity
Innovation

Methods, ideas, or system contributions that make the work stand out.

synthetic clinical benchmarks
utility-constrained realism
missingness structure
care-gap benchmark
source fidelity
🔎 Similar Papers
No similar papers found.