Learning Latent Protein Languages for Autoregressive Generation

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limited performance of autoregressive Transformers in protein generation and the lack of contextual semantics in conventional discrete representations by proposing two learned latent protein languages: PLL and SLL. Methodologically, we construct an ESM-2-based contextual alphabet and an improved GCP-VQVAE structural tokenizer, training autoregressive models via next-token prediction. Results demonstrate that PLLM reduces low-entropy samples by 54%, while SLL decreases validation perplexity by 34%, substantially optimizing computational scaling laws. Furthermore, long-protein generation achieves an approximately 1000-fold speedup over AlphaFold2, realizing concurrent breakthroughs in both generative quality and efficiency.
πŸ“ Abstract
Autoregressive transformers remain comparatively weak for protein sequence and structure generation. We study the role of target representation: amino acid tokens encode residue identities without explicit contextual semantics, while backbone coordinates require a discrete representation in our framework. We introduce two learned latent protein languages. Protein Latent Language (PLL) maps sequences to a 4,096-state contextual alphabet built on a frozen ESM-2 encoder, with one token per residue. Structure Latent Language (SLL) adapts GCP-VQVAE Lite with auxiliary sequence and confidence supervision while retaining decoding to backbone coordinates. We separately pretrain autoregressive transformer models on PLL and SLL tokens using next-token prediction, yielding PLLM and SLLM. Under matched downstream sequence training, PLLM has a fitted compute-scaling exponent of 0.038 versus 0.020 for the amino acid autoregressive model. In unconditional sequence generation, PLLM reduces the fraction of samples below a heuristic 1.5-bit residue-composition entropy threshold by 54% relative to the amino acid model across sampling temperatures. For sequence-to-structure prediction, replacing the original GCP-VQVAE Lite tokenizer with SLL reduces best validation perplexity by 34% under matched training. For long proteins, latent-token sampling is approximately 1,000 times faster than MSA-based AlphaFold2 in our measurements. In backbone generation, SLLM compares favorably with other generative models on diversity and novelty. We also observe early signs that using SLLM's internal token confidence for inference-time sampling can improve sequence-to-structure prediction quality beyond a single decoded sample. These results position learned latent protein languages as a promising substrate for autoregressive transformer scaling and inference-time sampling in protein generation.
Problem

Research questions and friction points this paper is trying to address.

autoregressive generation
protein sequence
protein structure
latent representation
transformer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Protein Language
Autoregressive Transformer
Protein Generation
Discrete Representation
Confidence-guided Sampling
M
Mahdi Pourmirzaei
University of Missouri
F
Farzaneh Esmaili
University of Missouri
A
Amir Ziashahabi
University of Southern California
M
Mohammadreza Pourmirzaei
Independent Researcher
Dong Xu
Dong Xu
Curators’ Distinguished Professor, University of Missouri
bioinformaticscomputational biologymachine learningartificial intelligencecomputer science