Distributional Balancing with Machine Learning for Clinical Trial Augmentation Using Real-World Data

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出DBML方法,通过匹配治疗组与外部数据库控制组的分布来解决临床试验中对照组难以平衡的问题,使用变分自编码器和重加权技术选择合适的对照样本。
📝 Abstract
In clinical trials, randomization of treatment and control groups is typically used to ensure the groups have similar covariate distributions on average, resulting in unbiased causal effect estimation. Such balanced covariate distributions are hard to achieve in practice, however, due to recruitment costs, patient dropouts, and more. One possible solution to this problem is to include control patients from external, real world databases. In this paper, we propose DBML, a method that selects control units from a real world database by matching the distribution between the treatment group and the potential control group, instead of matching units between two groups. DBML has three steps. First, we detect anomalous database units (with respect to the treatment distribution) using a variational autoencoder. Second, we re-weight the remaining database units to match the distribution of the treatments. Finally, we use these weights to sample units for our control group. The proposed method is compared to alternative matching-based algorithms and weighted algorithms, achieving superior performance in covariate balance.
Problem

Research questions and friction points this paper is trying to address.

clinical trials
covariate balance
real-world data
Innovation

Methods, ideas, or system contributions that make the work stand out.

DBML
distributional balancing
variational autoencoder
real-world data
💼 Related Jobs
No related jobs found.
Z
Zern Ke
Department of Statistics, Rutgers University, New Brunswick
M
Mingshi Cui
Department of Statistics, Rutgers University, New Brunswick
Gemma Moran
Gemma Moran
Assistant Professor of Statistics, Rutgers University
Javier Cabrera
Javier Cabrera
Professor of Statistics, Cardiovascular Institute, Rutgers University
StatisticsBiostatisticsAI/MLData miningbig data