Score
Design, build, or analyze generative modeling systems that produce peptide sequences and corresponding 3D structures conditioned on target properties or labels, including support for noncanonical residues and terminal chemical modifications. Implement sequence–structure co-design and full-atom peptide design using conditional architectures (e.g., conditional GANs, VAEs with latent diffusion and equivariant backbones), latent-space sampling and denoising, and block-level latent representations to generate diverse, physically plausible candidate peptides.
This study systematically compares the heterogeneous challenges in generative modeling of small-molecule and therapeutic peptide drugs using diffusion models: small molecules prioritize synthetic accessibility, whereas peptides require concurrent optimization of biostability, correct folding, and low immunogenicity; both domains suffer from inaccurate scoring functions, scarcity of high-quality experimental data, and absence of experimental validation loops. Method: We propose a unified iterative denoising framework, customized with graph-based representations for small molecules and joint sequence-structure embeddings for peptides, integrated with physicochemical property optimization, stability modeling, and immunogenicity assessment. Results: Experiments demonstrate significant improvements in molecular diversity and target-binding accuracy; however, critical bottlenecks persist in synthetic feasibility and data quality. Contribution: This work is the first to characterize fundamental design disparities between these drug modalities under a unified diffusion paradigm and establishes an experimentally grounded closed-loop optimization pathway.
Deep learning in protein design faces two core challenges: inadequate modeling of sequence–structure–function relationships and poor out-of-distribution generalization. To address these, we propose the first architecture-centric taxonomy of protein design models, systematically categorizing methodological evolution—from unimodal structure prediction (e.g., AlphaFold, ESMFold) and generative sequence design (e.g., ProteinMPNN, RFdiffusion) to joint sequence–structure–function co-design (e.g., ESM3). We identify generalizable fitness landscape modeling as the critical pathway to overcoming generalization bottlenecks. By synthesizing advances in attention mechanisms, diffusion modeling, multimodal representation learning, and geometric deep learning, we delineate the capability boundaries of state-of-the-art models and establish joint co-design as the optimal paradigm. This framework provides a systematic methodology for transcending natural evolutionary constraints and enabling rational design of novel functional proteins.
Current research in protein generative modeling suffers from significant fragmentation in representation schemes, model architectures, and task formulations, lacking a unified evaluation framework. This work presents the first comprehensive framework that systematically integrates sequence-based, geometric, and multimodal representations by unifying SE(3)-equivariant diffusion, flow matching, and hybrid prediction-generation architectures. We introduce a data partitioning strategy that prevents information leakage, a physics-informed validation mechanism for structural plausibility, and a function-oriented evaluation protocol. Furthermore, we establish a systematic taxonomy and benchmark spanning tasks from structure prediction to protein–protein interactions. Our study provides both a methodological foundation and practical guidelines for reliable, function-driven protein engineering, while highlighting critical challenges such as conformational dynamics modeling and biosafety considerations.
Antimicrobial peptide (AMP) design faces challenges stemming from the vast sequence space, poor interpretability of generative model latent spaces, and low optimization efficiency. To address these, we propose a physicochemical-property-guided dimensionality reduction framework: first, a variational autoencoder learns a latent representation of AMPs; second, principal component analysis (PCA) compresses this latent space into an interpretable low-dimensional subspace guided by key physicochemical properties—such as hydrophobicity and net charge; finally, within this compressed space, semi-supervised learning and Bayesian optimization jointly enable efficient, activity-directed AMP discovery under limited labeled data. Experiments demonstrate that our method improves search efficiency by 2.3× over baselines while enhancing latent-space interpretability and maintaining high generation fidelity for potent AMPs. This work establishes a new rational design paradigm for functional peptides in data-scarce regimes.
Optimizing antibody complementarity-determining regions (CDRs) for developability faces challenges of low search efficiency in the raw sequence space and high evaluation costs due to black-box, non-differentiable metrics (e.g., aggregation propensity, expression yield). To address this, we propose LEAD—a deep generative framework that learns a shared latent space jointly encoding sequence and structure, enabling their co-optimization. Crucially, LEAD introduces a gradient-free black-box guidance strategy, allowing efficient optimization with respect to arbitrary, non-differentiable developability objectives. In both single- and multi-objective CDR design tasks, LEAD reduces query count by over 50% compared to state-of-the-art baselines, while yielding higher-quality candidates. This work establishes a scalable, high-fidelity paradigm for joint sequence–structure antibody design, advancing computational antibody engineering.
This work addresses the challenge of efficiently generating target-specific peptides with coordinated sequence and structure design under full-atom geometric constraints. To this end, the authors propose MEET, a memory-efficient E(3)-equivariant Transformer that achieves linear memory scaling during encoding, decoding, and denoising by jointly propagating scalar and vector feature streams. The architecture innovatively reformulates geometric computations into a memory-efficient attention mechanism, initializing vector features via global coordinate aggregation, incorporating distance-augmented dot products, and injecting covalent bond constraints through sparse key adaptation. Integrated within a variational autoencoder and latent diffusion framework, MEET significantly enhances the binding affinity, physical plausibility, and diversity of generated peptides on the large-scale AFDB dataset, outperforming current state-of-the-art methods.
Existing generative models struggle to directly model all-atom protein structures due to the variable-length nature of side chains, which complicates joint sequence–structure modeling. To address this, we propose a partially implicit protein representation: the backbone geometry is explicitly modeled, while side-chain conformations and amino acid identities are jointly encoded as fixed-dimensional residue-level latent variables. We employ flow matching to jointly model the full-atom structure and sequence distributions in this latent space. Our method integrates geometric deep learning and equivariant neural networks to enable 3D-structure-aware generation. It is the first approach to efficiently generate proteins up to 800 residues long with high structural validity and controllable sequence design. It outperforms state-of-the-art methods across atomic-level co-designability, conformational diversity, structural validity, and motif-scaffolding tasks—demonstrating scalability, robustness, and practical utility for de novo protein design.
本文提出SimpleDesign模型,通过单阶段端到端训练直接在数据空间中联合设计蛋白质序列和结构,解决了多模态关系生成模型的复杂训练问题。
This work proposes the first single-stage, all-atom protein co-design framework that overcomes the limitations of traditional two-stage approaches, which struggle to achieve atomic-level precision and flexibly incorporate non-canonical amino acids. Built upon a unified multimodal diffusion model, the method jointly generates discrete atom types and continuous atomic coordinates in a single forward pass, directly inferring residue identities from the resulting atomic arrangements. This paradigm inherently supports non-canonical amino acids and eliminates the artificial separation between sequence and structure design. Evaluated on both unconditional protein generation and protein binder design tasks, the approach substantially outperforms existing single- and two-stage methods, achieving up to a tenfold increase in success rate on challenging design benchmarks.
Existing protein generative models typically require predefined sequence lengths, limiting design flexibility. This work proposes the Generalized Poisson Flow (GPFlow) framework, which introduces a non-homogeneous generalized Poisson process into protein generation for the first time. By learning the associated rate function, GPFlow enables joint modeling of sequences and structures without requiring fixed-length inputs. The method unifies Euclidean, categorical, and Riemannian geometric modalities, supporting unconditional design, motif scaffolding, and peptide co-design tasks, with theoretical guarantees for target distribution recovery. Experiments demonstrate that GPFlow accurately reproduces natural length distributions, achieves state-of-the-art performance in 10 out of 16 motif scaffolding tasks, and shows competitive results in peptide co-design, significantly outperforming fixed-length baselines.
该研究通过引入集合条件引导框架,优化分子构象集合的模式和属性,解决分子设计中考虑单一生物活性构象的问题。
This work addresses the unreliability of optimization directions in reinforcement learning–driven cyclic peptide generation, which often arises when permeability predictors are extrapolated beyond their domain of applicability. To mitigate this issue, the study introduces conformal prediction into molecular generative design for the first time, constructing an uncertainty-aware permeability predictor integrated within a reinforcement learning framework such as PepINVENT. By dynamically constraining the generation process to high-confidence chemical regions at user-specified confidence levels, the proposed method substantially reduces futile exploration and enhances both the reliability and efficiency of designing highly permeable cyclic peptides. Furthermore, it provides statistically calibrated uncertainty guarantees for the predicted properties of candidate molecules.