Score
Identifying biologically meaningful sequence or circuit motifs from model reconstructions by validating that recovered circuits reproduce the generative distribution and by representing cis-regulatory control so the model predicts how perturbations change transcriptional regulation.
This study addresses the challenging problem of predicting transcriptional responses to unseen gene perturbations in single cells under a zero-shot setting, where perturbation effects arise from complex interactions among cellular states, gene functions, and regulatory mechanisms. To tackle this, the authors propose CisTransCell, a novel framework that, for the first time, incorporates both cis-regulatory sequences and trans-encoded protein sequences as multimodal priors alongside cellular expression states. CisTransCell models the cascade from gene function through regulatory logic to downstream transcriptional changes. By integrating genomic, proteomic coding, and single-cell transcriptomic data, the method substantially outperforms existing approaches across multiple benchmark datasets, demonstrating exceptional performance in zero-shot gene perturbation prediction.
This study addresses the challenge of spurious edge generation when reconstructing causal regulatory networks of dynamical systems from discrete-state trajectories. The authors propose a modeling framework based on integral-form additive nonparametric ordinary differential equations, which accommodates both dense and sparse irregular sampling. By incorporating a data-driven edge selection mechanism, the method effectively suppresses false connections while inferring time-varying, weighted, bidirectional causal relationships between nodes, including activation or inhibition effects. This enables the construction of interpretable, symbolic causal networks. In five simulation experiments, the approach substantially outperforms GRADE, reducing spurious edges from 239 to zero in the most challenging scenario while nearly preserving all true regulatory links, thereby achieving high-precision dynamic network reconstruction.
In genomic foundation models, sequence predictability is often confounded with genuine regulatory signals, impeding the interpretation of non-coding regulatory function. This study introduces a residualization and permutation diagnostic framework, integrated with in silico mutagenesis (ISM), to clearly disentangle sequence prediction layers from regulatory output layers across three architecturally distinct models—Caduceus-Ph, HyenaDNA, and Enformer—revealing no overlap between the two at top regulatory elements. The work demonstrates that proximal regulatory boundaries are robustly confined within 10 kb, that a simple six-feature linear model suffices to recapitulate classification of the top 10% Caduceus-predicted elements (AUC = 0.985), and that top elements common to all three models are significantly enriched for brain eQTLs (3.3-fold enrichment, p < 5 × 10⁻³).
This work addresses the challenge in single-cell perturbation prediction that perturbation-specific signals are sparse and often obscured by dominant invariant expression structures, hindering existing methods from learning generalizable causal representations. To overcome this, the authors propose PerturbedVAE, a novel framework that explicitly disentangles perturbation-specific and invariant features for the first time, guided by identifiability theory to recover sparse perturbation effects. Built upon a variational autoencoder architecture, PerturbedVAE supports modeling of combinatorial perturbations and achieves substantially improved out-of-distribution prediction performance on standard benchmarks. Furthermore, the model reveals interpretable perturbation–response mechanisms, offering new insights into cellular response dynamics.
This work addresses the lack of interpretability in autoregressive protein language models, which obscures their cross-layer computational dynamics. To this end, we propose ProGenMech, a framework that introduces cross-layer transcoders (CLTs) into autoregressive protein generation for the first time, enabling mechanistic interpretation of ProGen3’s generative and functional prediction processes through reconstruction of sparse latent variables. By integrating sparse mixture-of-experts (MoE) architectures with zero-shot circuit discovery, ProGenMech effectively identifies interpretable latent circuits associated with protein function. Experimental results demonstrate that ProGenMech outperforms baseline methods in causal generation and zero-shot fitness prediction tasks, accurately reproduces the original model’s output distribution, and successfully localizes biologically meaningful functional motifs and evolutionarily conserved sequence regions.
This study addresses the limitation of existing gene regulatory network inference methods, which predominantly focus on pairwise interactions and struggle to accurately identify sets of co-regulators. To overcome this, the authors propose BRIDGE, a novel framework that enables end-to-end inference of complete regulatory sets for the first time, accompanied by the TRACE diagnostic suite to pinpoint pipeline bottlenecks. Key innovations include a mechanism-mismatch stress test designed to prevent information leakage and circular dependencies, and Residual HOS²—a residual higher-order set scoring method that directly models raw expression vectors without handcrafted features. The approach integrates PairS² for candidate generation with HOS²-based reranking. Evaluated across 30 cooperative regulatory settings, BRIDGE achieves a Jaccard similarity of 0.460, recall of 0.597, and doubles the exact recovery rate to 0.113, while reducing candidate set size by 94–97% without compromising performance.
This study addresses the limitations of existing methods in predicting cis-regulatory element (CRE) activity from DNA sequences, which often suffer from insufficient accuracy and lack of interpretability. The authors propose R3LM, a novel framework that introduces CRE-ReasonBench—the first dataset incorporating mechanistic reasoning trajectories—and employs a two-stage training strategy to integrate structured biological knowledge into a large language model (LLM) for reasoning-based regression. Evaluated across three cell types, R3LM significantly outperforms both sequence-only LLMs and specialized DNA models, achieving state-of-the-art performance in enhancer activity prediction while simultaneously providing interpretable insights into the underlying regulatory mechanisms. This approach thus bridges high predictive accuracy with biologically meaningful interpretability.
This work proposes a novel automated algorithm for neural circuit discovery grounded in formal neural network verification, addressing the limitations of existing heuristic-based approaches that lack provable guarantees over continuous input domains. For the first time, the method provides three formally verifiable properties: input-domain robustness, robust reparability, and minimality, while elucidating their intrinsic theoretical connections. By integrating state-of-the-art neural verifiers with formal reasoning techniques, the approach enables provably correct circuit extraction. Experimental results demonstrate that circuits extracted from multiple vision models using this framework significantly outperform those obtained via conventional methods in terms of robustness, thereby establishing a rigorous theoretical foundation for provable circuit discovery in deep neural networks.
This work proposes MechPert, a lightweight framework designed to predict transcriptional responses to unobserved genetic perturbations and guide experimental design. Leveraging a large language model–driven multi-agent system, MechPert independently generates directed regulatory hypotheses with associated confidence scores, then aggregates these through mechanistic consensus to filter spurious associations and construct a weighted neighborhood for downstream prediction. Unlike conventional approaches that rely on functional similarity or static knowledge graphs, MechPert introduces mechanistic consensus as an inductive bias, prioritizing directed regulatory logic over symmetric co-occurrence relationships. Evaluated on four Perturb-seq cell line benchmarks, MechPert achieves up to a 10.5% improvement in Pearson correlation coefficient under low-data regimes (N=50) and demonstrates up to a 46% gain in performance for anchor gene selection in experimental design compared to traditional network centrality methods.
This study addresses the limited interpretability of internal representations in genomic language models, which often conflate genuine biological signals with sequence artifacts. The authors propose a novel approach combining sparse dictionary learning with causal interventions to extract interpretable features from the hidden activations of Nucleotide Transformer and DNABERT-2. To mitigate confounding effects from GC content and repetitive elements, they introduce a composition-matched, site-resolved validation protocol. Through directional ablation experiments, they establish—for the first time—the causal role of specific model features in mediating cell type–specific transcription factor binding (CTCF, GATA1, REST) during forward propagation. The method consistently recovers 7–14 causal features per factor, with no signal detected in negative controls, demonstrating both reliability and reproducibility.