PFArena: Benchmarking Language Models for Protein Modification

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of standardized evaluation benchmarks and the high cost of experimental validation in protein modification by constructing an open-source benchmark encompassing four distinct scenarios. Methodologically, it introduces a novel controlled task interface reflecting varying degrees of prior data availability to enable fair comparisons across multiple model families. By integrating protein language models (PLMs), large language models (LLMs), and agent-based techniques alongside complementary metrics, this work systematically evaluates single-mutation generation and multi-mutation ranking performance. The findings reveal how model efficacy evolves with data availability, demonstrating that PLMs excel at mutation generation while LLMs outperform in ranking tasks. Ultimately, this research establishes a reliable evaluation framework for protein engineering.
📝 Abstract
Protein modification requires navigating an immense sequence space, yet wet-lab validation remains low-throughput and costly. Although computational paradigms including protein language models (PLMs), large language models (LLMs), and LLM-based agents have shown promise in protein modification, their relative efficacy across realistic experimental decision-making settings remains unclear. To bridge this gap, we introduce PFArena, a benchmark comprising four controlled task interfaces that cover single-mutant generation and multi-mutant ranking. By providing varying levels of mutation fitness data, PFArena reflects four representative research scenarios characterized by differing degrees of prior experimental context. We assess six PLMs, six LLMs, and five LLM-based agents using complementary metrics to measure both peak and overall protein modification performance. Our evaluation reveals that model performance shifts systematically with the availability of target-specific experimental evidence: PLMs demonstrate proficiency in open-ended single-mutant generation by leveraging protein-specific priors, whereas LLMs and agents perform strongly in multi-mutant ranking, particularly when target-specific fitness data are available. Nevertheless, all model families face fundamental challenges with increasing search-space size and mutation depth. We release our code and benchmark suite to facilitate reproducible research in model-assisted protein modification.
Problem

Research questions and friction points this paper is trying to address.

Protein modification
Benchmarking
Protein language models
Large language models
LLM-based agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Protein modification
Benchmark
Protein language models
Large language models
LLM-based agents
🔎 Similar Papers
No similar papers found.