Nonlinear multi-study factor analysis

📅 2026-01-26
📈 Citations: 0
Influential: 0
📄 PDF

career value

204K/year
🤖 AI Summary
Disentangling shared and study-specific latent factors from high-dimensional data across multiple studies remains a key challenge in cross-disease gene expression analysis. This work proposes a nonlinear multi-study sparse variational autoencoder that, for the first time, integrates a sparse nonlinear factor model into a multi-study framework. By implicitly penalizing the number of latent factors and modeling sparse dependencies between features and latent factors, the method automatically separates shared from study-specific components. Theoretical analysis establishes identifiability of the latent factors, and experiments on platelet gene expression data successfully recover biologically meaningful co-expression modules, demonstrating both the effectiveness and interpretability of the approach.

Technology Category

Application Category

📝 Abstract
High-dimensional data often exhibit variation that can be captured by lower dimensional factors. For high-dimensional data from multiple studies or environments, one goal is to understand which underlying factors are common to all studies, and which factors are study or environment-specific. As a particular example, we consider platelet gene expression data from patients in different disease groups. In this data, factors correspond to clusters of genes which are co-expressed; we may expect some clusters (or biological pathways) to be active for all diseases, while some clusters are only active for a specific disease. To learn these factors, we consider a nonlinear multi-study factor model, which allows for both shared and specific factors. To fit this model, we propose a multi-study sparse variational autoencoder. The underlying model is sparse in that each observed feature (i.e. each dimension of the data) depends on a small subset of the latent factors. In the genomics example, this means each gene is active in only a few biological processes. Further, the model implicitly induces a penalty on the number of latent factors, which helps separate the shared factors from the group-specific factors. We prove that the latent factors are identified, and demonstrate our method recovers meaningful factors in the platelet gene expression data.
Problem

Research questions and friction points this paper is trying to address.

multi-study factor analysis
shared and specific factors
high-dimensional data
nonlinear factor model
latent factor identification
Innovation

Methods, ideas, or system contributions that make the work stand out.

nonlinear multi-study factor analysis
sparse variational autoencoder
shared and study-specific factors
latent factor identifiability
high-dimensional genomics