A Goemans-Williamson type algorithm for identifying subcohorts in clinical trials

📅 2025-06-12
📈 Citations: 0
Influential: 0
📄 PDF

career value

193K/year
🤖 AI Summary
High patient heterogeneity in clinical trials impedes identification of biologically homogeneous subpopulations. To address this, we propose a semidefinite programming (SDP)-based subcohort discovery framework incorporating biologically interpretable constraints. Methodologically, we introduce the first SDP relaxation that integrates methylome–transcriptome co-variation priors and design a Goemans–Williamson-type randomized rounding algorithm with a theoretical approximation ratio of 0.82. Applied to the Curtis breast cancer cohort, our method identifies a clinically meaningful subcohort enriched for metastatic cases and systematically uncovers an interpretable subcohort characterized by coordinated hypermethylation of tumor suppressor genes and altered nuclear receptor expression. This work establishes a computationally rigorous, biologically grounded, and experimentally verifiable paradigm for mechanistic disease dissection and targeted therapeutic development.

Technology Category

Application Category

📝 Abstract
We design an efficient algorithm that outputs a linear classifier for identifying homogeneous subsets (equivalently subcohorts) from large inhomogeneous datasets. Our theoretical contribution is a rounding technique, similar to that of Goemans and Williamson (1994), that approximates the optimal solution of the underlying optimization problem within a factor of $0.82$. As an application, we use our algorithm to design a simple test that can identify homogeneous subcohorts of patients, that are mainly comprised of metastatic cases, from the RNA microarray dataset for breast cancer by Curtis et al. (2012). Furthermore, we also use the test output by the algorithm to systematically identify subcohorts of patients in which statistically significant changes in methylation levels of tumor suppressor genes co-occur with statistically significant changes in nuclear receptor expression. Identifying such homogeneous subcohorts of patients can be useful for the discovery of disease pathways and therapeutics, specific to the subcohort.
Problem

Research questions and friction points this paper is trying to address.

Designing an efficient algorithm for identifying homogeneous subcohorts in clinical data
Approximating optimal solutions for inhomogeneous dataset classification
Identifying patient subcohorts with significant biomarker changes for targeted therapy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Goemans-Williamson rounding for subcohort identification
Linear classifier for homogeneous patient subsets
Statistical test for methylation-expression co-occurrence