Score
Designs, implements, and evaluates variational autoencoder models (including hybrid VAEs) that encode sequence-structured polymers—e.g., random heteropolymer (RHP) sequences—into continuous latent spaces and decode latent vectors to reconstruct or generate new sequences and monomer compositions. Builds tooling to capture and disentangle sequence and chemical modalities in the latent space and to analyze or sample that space to propose candidate sequences and compositions.
This work addresses the lack of efficient computational tools for guiding the design of random heteropolymers (RHPs) to mimic protein functions. To this end, the authors propose DeepRHP, a semi-supervised hybrid variational autoencoder that, for the first time, jointly embeds chemical features and sequence information into a unified latent space at both feature and sequence levels, enabling multifunctional, structure-constrained generative design. The framework flexibly incorporates arbitrary relevant attributes, significantly enhancing controllable RHP sequence generation. Experimental results demonstrate that DeepRHP successfully predicts RHP monomer compositions capable of stabilizing membrane proteins such as Aquaporin Z, with predictions in strong agreement with published experimental data, thereby validating its effectiveness and practical utility.
In Bayesian optimization (BO) over complex structured spaces—such as molecular spaces—conventional variational autoencoder (VAE)–Gaussian process (GP) coupling leads to poor generalization and necessitates task-specific architectural customization due to strong interdependence between the VAE’s latent space and the surrogate model. To address this, we propose a decoupled BO framework that independently trains the VAE generative model and the GP surrogate, coordinating them via Bayesian updating: the VAE learns transferable structural representations, while the GP models uncertainty of the objective function in the fixed latent space. This decoupling eliminates reliance on task-specific latent geometries, enhancing both optimization efficiency and stability under limited evaluation budgets. On multiple molecular optimization benchmarks, our method identifies higher-performing candidate molecules with fewer evaluations than state-of-the-art coupled VAE–GP approaches.
Biological sequence design requires optimizing high-dimensional continuous functional properties—such as fluorescence spectra, photostability, or antimicrobial activity—of DNA, RNA, or peptides. However, existing models are limited to binary labels and fail to capture the complex geometry of such property manifolds. To address this, we propose PrIVAE: a graph neural network–based variational autoencoder incorporating isometric regularization and nearest-neighbor graph constraints to rigorously preserve the intrinsic geometric structure of the property manifold in latent space. PrIVAE enables differentiable inverse design for high-dimensional continuous properties—the first method to achieve this capability. It maintains high sequence reconstruction fidelity while substantially improving the generation efficiency of rare functional sequences (e.g., DNA nanoclusters with target emission wavelengths). Wet-lab validation demonstrates a 16.1× enrichment in desired variants, establishing a new paradigm for function-driven, rational biological sequence design.
To address the decoupling of prediction and generation capabilities in generative models for scientific discovery, this paper proposes a VAE-DKL synergistic framework: deep kernel learning (DKL) is embedded into the latent space of a variational autoencoder (VAE), enabling structured modeling and property-guided optimization of latent variables via Gaussian process (GP) regression. This approach achieves, for the first time, end-to-end joint training of generation and property prediction. On the QM9 dataset, it attains a prediction error of 0.12 eV for enthalpy—surpassing a standalone VAE+GP baseline—and enables controllable generation of novel molecular structures outside the training set that satisfy target property constraints. The core innovation lies in the co-optimization mechanism within the latent space: it preserves the VAE’s high-fidelity sampling capability while endowing it with interpretable and intervenable property prediction functionality, thereby significantly enhancing efficiency in materials inverse design.
This paper addresses the fundamental mismatch between the continuous latent space of standard Variational Autoencoders (VAEs) and the inherently discrete nature of data such as text. To resolve this, we propose the Discrete VAE—a VAE explicitly designed for categorical latent variables. Methodologically, we derive the evidence lower bound (ELBO) rigorously from first principles of variational inference under categorical latents and employ the Gumbel-Softmax reparameterization to enable differentiable gradient estimation in discrete latent spaces. Our key contributions are threefold: (1) a tutorial-style, unified theoretical framework for discrete VAEs; (2) a robust and reproducible training paradigm; and (3) publicly released, fully functional code. Experiments demonstrate that the Discrete VAE significantly improves interpretability and structural coherence in discrete data generation, outperforming continuous-latent baselines while preserving principled probabilistic modeling.
This study addresses three key challenges in molecular generation: broad chemical space coverage, weak property controllability, and low generation efficiency. Methodologically, we propose a Transformer-based variational autoencoder (VAE) framework that employs SELFIES for syntax-valid molecular representation; adopts an encoder–decoder latent variable architecture to jointly model unconditional and property-conditioned generation; and integrates LoRA (Low-Rank Adaptation) for efficient fine-tuning. Evaluated on the GuacaMol and MOSES benchmarks, our approach achieves state-of-the-art (SOTA) performance or surpasses leading baselines across multiple metrics. In the Tartarus docking task, it significantly improves the distribution of predicted binding affinities, demonstrating superior property precision control, structural diversity, and chemical validity. Collectively, these results validate the framework’s effectiveness in balancing broad chemical space exploration with targeted property optimization and high-fidelity generation.
This work addresses the challenge of integrating variational autoencoders (VAEs) as trainable layers within neural networks. It proposes a general framework for flexibly embedding VAEs into arbitrary network architectures, accompanied by an end-to-end training strategy that leverages the reparameterization trick and probabilistic modeling to ensure full differentiability throughout the pipeline. For the first time, this approach enables VAEs to function as plug-and-play modules akin to standard neural network layers, substantially enhancing their compatibility and representational capacity within complex models. Experimental results demonstrate that the proposed VAE layer consistently achieves stable performance across diverse tasks and outperforms conventional standalone VAE models, thereby significantly expanding the applicability of VAEs in deep learning systems.
本文提出SimpleDesign模型,通过单阶段端到端训练直接在数据空间中联合设计蛋白质序列和结构,解决了多模态关系生成模型的复杂训练问题。
Standard variational autoencoders employ Gaussian priors, which struggle to align with data manifolds exhibiting non-Euclidean topologies—such as periodicity or boundedness—leading to distorted representations. This work proposes a topology-aware latent space modeling framework that constructs factorized prior distributions tailored to manifolds decomposable into products of circles, intervals, and lines, along with their finite group quotients. This design enables disentangled latent representations and analytically tractable KL divergences. By integrating differentiable coordinate transformations, group-invariant decoding, and anchor-point constraints, the approach ensures smooth gradients and topological consistency. To our knowledge, this is the first method to systematically align latent variable distributions with the intrinsic topology of data manifolds, supporting reparameterizable encoder–prior pairs and significantly outperforming Gaussian-prior baselines on synthetic manifolds as well as rotation- and cyclic-translation variants of MNIST.
In molecular generative models, the SELFIES string representation often introduces artifacts—such as those related to sequence length, branching, and ring structures—that obscure genuine chemical signals in the latent space. This work addresses this challenge by employing linear probes within a frozen Transformer-VAE latent space to identify global directions governing key physicochemical properties. The authors propose a confounding-aware evaluation framework that combines residualization analysis with decoded molecular traversal to systematically disentangle representation artifacts from true chemical signals. Using this approach, they demonstrate for the first time that six properties—including cLogP and FractionCSP3—exhibit robust and monotonically controllable directions even within an entangled latent space, thereby validating the feasibility of chemically meaningful latent space editing.