Deep and Probabilistic Models for Gene Regulatory Network Inference

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses key limitations in gene regulatory network (GRN) reconstruction—namely, the tight coupling between modeling assumptions and inference methods, the lack of uncertainty quantification, and the difficulty of transferring prior knowledge across species. To overcome these challenges, the authors propose a two-stage framework: first, a transferable prior model, GLM-Prior, is built by fine-tuning the Nucleotide Transformer on DNA sequences to predict transcription factor–target interactions; second, a probabilistic matrix factorization model, PMF-GRN, incorporates variational inference to enable uncertainty-aware network refinement. This approach uniquely integrates transferable deep sequence modeling with probabilistic inference that quantifies uncertainty, significantly improving the accuracy, generalizability, and reliability of GRN reconstruction across yeast, mouse, and human datasets, while enabling robust evaluation even when reference networks are incomplete.
📝 Abstract
Gene regulatory networks (GRNs) link transcription factor (TF) proteins to their target genes, yet reconstructing these networks from genome-wide data remains challenging under practical and methodological constraints. Many methods couple modeling assumptions to a specific inference procedure and rely on heuristic model selection, while evaluation is constrained by incomplete reference networks and point-estimate outputs that lack uncertainty. GRN reconstruction also depends on prior knowledge to constrain TF-gene interactions, yet available priors are often assay-dependent and difficult to transfer across species and less-characterized systems. In this thesis, we develop two complementary frameworks that address these limitations. In the first, PMF-GRN casts GRN inference as a probabilistic graphical model optimized by variational inference, enabling principled model selection and uncertainty-aware edge estimates. In the second, GLM-Prior addresses the prior bottleneck by fine-tuning the pretrained Nucleotide Transformer to predict TF-target gene interactions directly from nucleotide sequence, while generalizing across yeast, mouse, and human settings. Together, PMF-GRN and GLM-Prior motivate a dual-stage view of GRN reconstruction in which sequence-derived priors provide a transferable starting scaffold and probabilistic inference refines regulatory estimates with quantified uncertainty under incomplete evaluation resources.
Problem

Research questions and friction points this paper is trying to address.

gene regulatory network
prior knowledge
uncertainty quantification
cross-species generalization
network inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

probabilistic graphical model
variational inference
Nucleotide Transformer
transferable prior
uncertainty quantification