A General Harness for Protein Foundation Model Fitness Prediction

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the distortion in fitness prediction caused by training data biases and input noise in protein foundation models. To this end, we propose VRH, a general, training-free, retrieval-augmented framework for mutation effect prediction. By integrating multiple sequence alignment (MSA) evidence with structural solvent accessibility information, VRH reliably calibrates model scores through uncertainty weighting and a gated background correction algorithm. Experimental results demonstrate that VRH improves the Spearman correlation coefficient by an average of 0.073 across three benchmark datasets. Furthermore, its derived model, VenusREM2, establishes new state-of-the-art performance records on multiple tasks. This work presents an efficient, fine-tuning-free paradigm for protein mutation effect prediction.
📝 Abstract
Accurate fitness prediction is central to protein engineering and understanding sequence-function relationships. With advances in deep learning, protein foundation models (PFMs) have become widely used for this task. Recent analyses, however, show that these models share preferences reflecting their training corpora, while unreliable inputs can further distort fitness predictions. Family-specific evolutionary evidence and structural context can help address these limitations by providing complementary constraints on model scores, motivating VenusREM-Harness (VRH), a general, model-agnostic, training-free Retrieval-Enhanced Mutation harness. It fuses frozen model scores with multiple sequence alignment (MSA) evidence according to model uncertainty, then applies gated background correction and score shrinkage based on structural confidence and solvent exposure. Across 1,211 assays and 3.1 million measured variants from ProteinGym, VenusMutHub, and the newly curated viral benchmark VenusViroHub, all 71 configurations improve Spearman correlation on all 3 benchmarks by 0.073 on average, with broad gains across 5 metrics. Extended analyses relate retrieval gains to model-MSA preference differences, assess domain-level gains and immune-escape cases, and quantify computational speedups. Built with VRH, VenusREM2 is the first to rank highest in all function, taxon, MSA-depth, and mutation-depth categories, with a ProteinGym Average Spearman of 0.556, 0.038 above the prior best.
Problem

Research questions and friction points this paper is trying to address.

protein fitness prediction
protein foundation models
sequence-function relationships
model bias
reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Protein Foundation Model
Retrieval-Enhanced Mutation
Training-free
Multiple Sequence Alignment
Fitness Prediction
🔎 Similar Papers
No similar papers found.