Score
Reconstructing gene regulatory networks from snapshot or time-series expression data while preserving gene identities across time and integrating mechanistic information (e.g., tumor expression and drug–target relationships) to produce patient-specific network models.
This study addresses the challenge of spurious edge generation when reconstructing causal regulatory networks of dynamical systems from discrete-state trajectories. The authors propose a modeling framework based on integral-form additive nonparametric ordinary differential equations, which accommodates both dense and sparse irregular sampling. By incorporating a data-driven edge selection mechanism, the method effectively suppresses false connections while inferring time-varying, weighted, bidirectional causal relationships between nodes, including activation or inhibition effects. This enables the construction of interpretable, symbolic causal networks. In five simulation experiments, the approach substantially outperforms GRADE, reducing spurious edges from 239 to zero in the most challenging scenario while nearly preserving all true regulatory links, thereby achieving high-precision dynamic network reconstruction.
This work addresses the limitation of existing biological foundation models, which predominantly rely on static gene expression data and thus fail to capture the dynamic evolution of gene regulatory networks (GRNs) during cellular development. The authors propose a novel approach that infers pseudotemporal trajectories from single-cell transcriptomic data, discretizes them into developmental snapshots, and reconstructs GRNs at each snapshot. A temporal graph neural network is then introduced to explicitly model the dynamic rewiring of regulatory interactions over time, enabling accurate prediction of gene expression, regulatory links, and key hub genes. To the best of our knowledge, this is the first application of temporal graph learning to single-cell biology. Evaluated on mouse erythroid gastrulation and pancreatic endocrine development datasets, the method significantly outperforms state-of-the-art foundation models such as scGPT and scFoundation across all three tasks, revealing non-trivial dynamic regulatory mechanisms.
This study addresses the challenge of disentangling molecular noise, cellular heterogeneity, and technical artifacts in single-cell gene expression data. We propose a Bayesian inference framework that integrates structured priors—temporal, spatial, or multimodal—with biochemically interpretable stochastic dynamical models. Methodologically, we introduce the first systematic unification of stochastic differential equation modeling, multimodal joint embedding, variational Bayesian inference, generative latent-variable models, and causal structure learning—overcoming the fundamental limitation that count-based data impose on mechanistic dynamical inference. Evaluated on both synthetic benchmarks and real spatiotemporal transcriptomic datasets, our approach enables high-accuracy estimation of gene-level regulatory parameters—including burst frequency, feedback strength, and microenvironmental response coefficients. It significantly enhances quantitative resolution of transcriptional feedback loops, RNA bursting kinetics, and tissue microenvironmental regulation. The framework establishes a novel, interpretable, and empirically verifiable paradigm for modeling regulatory networks in multicellular systems.
Accurately predicting individual drug response from pre-treatment transcriptomes remains challenging due to the scarcity of clinical response labels and post-treatment molecular data. This work proposes a novel framework that first constructs patient-specific gene regulatory networks integrated with drug target information, then leverages a LINCS L1000–pretrained gene attention model to simulate drug perturbation effects. A CLIP-style contrastive learning strategy aligns personalized knowledge graphs with perturbation representations in a shared latent space. The approach uniquely unifies mechanistic interpretability with dynamic perturbation modeling and enables zero-shot transfer. It consistently outperforms state-of-the-art methods across multiple TCGA partitioning schemes and achieves a 5.6% improvement in zero-shot AUROC on the I-SPY2 trial, while yielding stable, biologically interpretable attributions at the gene and pathway levels.
This study addresses the challenge of improving out-of-distribution (OOD) generalization for predicting transcriptional responses to genetic perturbations—specifically, unseen single- and double-gene perturbations and novel cell lines. To overcome the limited experimental coverage that constrains existing methods, we introduce, for the first time, multi-source biological knowledge graphs to guide OOD modeling, establishing a rigorous benchmark framework that enforces strict cross-perturbation-type and cross-cell-line generalization. Our method integrates graph neural networks for knowledge graph encoding, multi-relational heterogeneous graph aggregation, OOD-aware training, and an interpretable attention mechanism. Across all three OOD settings, TxPert achieves a mean R² improvement of 12.7% over state-of-the-art baselines. We publicly release both the new benchmark and the model implementation.
This study addresses the challenge of quantitatively modeling small-molecule–induced perturbations in cellular gene expression to elucidate drug transcriptional mechanisms, predict off-target effects, and identify drug repurposing opportunities. To overcome the limitation of existing knowledge graph (KG) models—restricted to binary drug–disease associations—we propose the first KG-driven graph neural network for gene expression perturbation prediction. Specifically, we construct a heterogeneous graph by integrating PrimeKG++ and LINCS L1000 data, incorporate multimodal embeddings from MolFormerXL and BioBERT, and design a graph attention network (GAT) coupled with a differential expression gene (DEG) prediction head to enable interpretable, cell-line– and compound-specific prediction of expression changes across 978 landmark genes. Ablation studies confirm that KG structural information substantially improves performance. Our model significantly outperforms MLP baselines under both scaffold and random splits, demonstrating robust generalization and enhanced biological interpretability.
This study investigates the feasibility of recovering causal relationships among genes from aggregated bulk gene expression data. By formalizing the notion of causal recoverability, the authors introduce two key criteria—functional form consistency and conditional independence consistency—and rigorously establish their necessary and sufficient conditions: accurate causal recovery is possible only when gene regulation follows linear aggregation and affine structural equation models. The theoretical analysis integrates causal inference with a functional consistency framework and is empirically validated on multiple bulk and single-cell datasets. Results demonstrate that real gene regulatory mechanisms are predominantly nonlinear, thereby revealing a fundamental limitation of current causal inference methods that rely solely on bulk expression data.
This study addresses the challenges of high dimensionality and temporal dependence in modeling dynamic co-expression networks from longitudinal omics data by proposing a novel approach based on multivariate linear mixed-effects models. The method captures both fixed and random effects of molecular features and leverages correlations among random effects to characterize node dependencies, thereby constructing time-evolving co-expression networks. Two innovative penalized algorithms grounded in thresholded covariance estimation are introduced to substantially improve the accuracy of network structure inference. Simulation studies demonstrate that the proposed method outperforms existing approaches in terms of mean squared error and mean absolute error. Applied to the CARDIA cohort data, it successfully uncovers temporal evolution patterns of protein co-expression networks and their associations with longitudinal trajectories.
This study addresses key limitations in gene regulatory network (GRN) reconstruction—namely, the tight coupling between modeling assumptions and inference methods, the lack of uncertainty quantification, and the difficulty of transferring prior knowledge across species. To overcome these challenges, the authors propose a two-stage framework: first, a transferable prior model, GLM-Prior, is built by fine-tuning the Nucleotide Transformer on DNA sequences to predict transcription factor–target interactions; second, a probabilistic matrix factorization model, PMF-GRN, incorporates variational inference to enable uncertainty-aware network refinement. This approach uniquely integrates transferable deep sequence modeling with probabilistic inference that quantifies uncertainty, significantly improving the accuracy, generalizability, and reliability of GRN reconstruction across yeast, mouse, and human datasets, while enabling robust evaluation even when reference networks are incomplete.
This study addresses the limited generalization of existing gene regulatory network inference methods under inductive settings and the absence of evaluation benchmarks aligned with biological experimental needs. The authors reformulate the task as a ranking-centric inductive graph completion problem and introduce the first co-evolutionary discrete diffusion framework that jointly models discrete gene expression states and regulatory interactions. To enhance training efficiency, they propose a TF-ALL subgraph sampling strategy. Furthermore, they establish BEELINE-KGC, the first inductive benchmark focused on discovering high-confidence novel regulatory relationships. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art models on BEELINE-KGC, and ablation studies confirm the effectiveness of each component.