Benchmarking Peptide-Protein Affinity Prediction Across Peptide and Target Shifts

๐Ÿ“… 2026-08-30
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
็ ”็ฉถ้€š่ฟ‡ๆ•ดๅˆ่‚ฝ-่›‹็™ฝ็ป“ๅˆๆ•ฐๆฎ๏ผŒ่ฏ„ไผฐไบ†ๅคš็ง่กจ็คบๆ–นๆณ•ๅ’Œๅ›žๅฝ’ๅ™จๅœจไธๅŒๆ•ฐๆฎๅˆ’ๅˆ†ไธ‹็š„ๆ€ง่ƒฝ๏ผŒไปฅ่งฃๅ†ณ่‚ฝ-่›‹็™ฝไบฒๅ’ŒๅŠ›้ข„ๆต‹็š„ๆณ›ๅŒ–้—ฎ้ข˜ใ€‚
๐Ÿ“ Abstract
Peptide-protein affinity models are often evaluated with a single data split, obscuring whether they interpolate among measurements for observed targets or generalize across peptide or target shifts. We integrated three sources of quantitative peptide-protein binding data to obtain 11,349 deduplicated pairs and benchmarked ten peptide representations, ESM-2 protein embeddings, and six regressors under peptide-similarity, within-target, and leave-target-out partitions. Across 60 matched representation-regressor configurations, mean test Spearman correlations were 0.462, 0.669, and 0.530, respectively. The top configuration shifted from ECFP-16 count fingerprints with random forest in the first two settings to HELM-BERT with Extra Trees when exact target sequences were excluded. Representation-rank correlations ranged from -0.042 to 0.624 across partitions, whereas regressor-rank correlations ranged from 0.771 to 0.943. Learning curves showed that representation differences were largest with limited supervision and narrowed as training data increased. PeptideCLM-2 adaptation and simple element-wise interaction features provided no consistent gain over a frozen encoder and direct concatenation under the tested protocols. These conclusions are specific to a dataset that pools transformed Kd, Ki, and IC50 measurements and to target exclusion at the exact-sequence level. Peptide-protein affinity benchmarks should therefore align data partitions with the intended use and jointly assess the effects of data scale, molecular representation, and downstream learner.
Problem

Research questions and friction points this paper is trying to address.

peptide-protein affinity
generalization
data partitioning
Innovation

Methods, ideas, or system contributions that make the work stand out.

peptide-protein affinity
data partitioning
molecular representation
regressors
generalization
J
Jiaxin Tian
College of Biology, Hunan University, Changsha, Hunan, China
D
Darren An
Lingang Laboratory, Shanghai, China
J
Jun Li
Lingang Laboratory, Shanghai, China