Linear Fitness Subspace in Protein Language Models Enables Sample-Efficient Directed Evolution

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the low sample efficiency caused by high-dimensional embeddings and the misalignment of zero-shot predictions in protein language models for directed evolution. We propose the linear fitness subspace hypothesis and introduce SGES, a method that reveals, for the first time, the existence of compact, experiment-specific linear directions within mutation-induced representation shifts. By estimating this subspace from few-shot data, SGES integrates Gaussian process regression with Bayesian optimization to enable surrogate modeling, uncertainty quantification, and active acquisition search. Evaluated on the ProteinGym benchmark, SGES significantly improves both fitness prediction accuracy and search efficiency, outperforming zero-shot models and mainstream baselines.
📝 Abstract
Model-guided directed evolution seeks to identify high-fitness protein variants under limited oracle budgets. Protein language models (PLMs) provide rich representations for this task, but task-agnostic zero-shot scores can be misaligned with a target assay, while supervised search in high-dimensional embedding spaces can make surrogate modeling and uncertainty estimation sample-inefficient. We propose the Linear Fitness Subspace (LFS) hypothesis: within mutation-induced residue-level representation changes, a compact, assay-specific set of directions makes fitness variation linearly accessible from few labeled variants. This is a local, supervision-recoverable statement rather than a claim that protein fitness landscapes or global PLM geometry are universally linear. Building on this observation, we introduce Subspace-Guided Evolutionary Search (SGES), which estimates an LFS from a small initial sample and performs surrogate modeling, uncertainty estimation, and acquisition in the learned subspace. Across 10 core ProteinGym assays, 87 extended static-validation assays, and an 18-assay budgeted-search evaluation, SGES improves fitness prediction and search efficiency over zero-shot PLMs and recent ML-guided protein optimization baselines. Controlled comparisons with PCA, random projections, label-shuffled PLS, classical mutation features, and acquisition ablations further isolate the benefit of a fitness-aligned site-delta coordinate.
Problem

Research questions and friction points this paper is trying to address.

directed evolution
protein language models
sample efficiency
fitness prediction
surrogate modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Protein Language Models
Directed Evolution
Linear Fitness Subspace
Subspace-Guided Evolutionary Search
Sample Efficiency
🔎 Similar Papers
2024-07-16arXiv.orgCitations: 0
S
SiYuan Ma
College of Computing and Data Science, Nanyang Technological University, Singapore
C
Canran Xiao
Sun Yat-sen University, China
Z
Zikai Xiao
Zhejiang University, China
A
Albert Gao
Carnegie Mellon University, USA
L
Liang He
Shanghai Institute of Optics and Fine Mechanics, China
X
Xuan-Yu Wang
Zhongnan Hospital, Wuhan University, China
S
Shuying Cao
University of Southern California, USA
Xiaojun Jia
Xiaojun Jia
Nanyang Technological University
Explainable AIRobust AIEfficient AI