🤖 AI Summary
This study addresses the underutilization of structural information and inefficiency in large-scale screening for protein binding prediction by proposing a dual-tower framework that integrates sequence, geometric, and topological features. Methodologically, persistent homology descriptors, including Vietoris-Rips filtrations and persistence landscapes, are introduced to characterize topological structures and combined with ESM-2 embeddings to construct a structure-aware Transformer. Cross-attention facilitates soft docking, while an NT-Xent contrastive loss enables joint optimization. Crucially, the architecture supports independent encoding to accelerate virtual screening. Experimental results demonstrate that the proposed model surpasses existing baselines on general protein-protein interaction benchmarks and is the only approach exceeding random performance on the PPB-Affinity task, significantly enhancing screening efficiency.
📝 Abstract
Protein binding prediction is central to target identification, therapeutic binder design, and large scale screening, yet remains challenging because binding depends on sequence, three dimensional geometry, and global structural organization. Recent folding models such as AlphaFold3 and Boltz-2 have substantially improved structure prediction, but their confidence outputs (pLDDT, pTM, ipTM) are not specifically designed for binary binding prediction, and dedicated structure aware predictors often require bound complex structures that are unavailable at screening scale. We introduce PIT-GCL, a dual tower structure aware framework that encodes each protein independently from its amino acid sequence, C{\alpha} point cloud, and a global persistent homology descriptor. Each tower combines residue ESM-2 embeddings with a topological summary computed from the H0 and H1 persistence landscapes of a Vietoris-Rips filtration, and processes the resulting tokens with a structure aware Transformer in which pairwise C{\alpha} distances enter as a learned attention bias. A bidirectional cross attention module then performs latent space soft docking between the two per-protein representations, and the model is trained with a combined binary cross entropy and NT-Xent contrastive objective. On three binary interaction prediction benchmarks, general PPI on PPIRef, TCRpMHC binding on STAG, and whole chain pairs on PPB-Affinity, PIT-GCL outperforms representative sequence based, structure aware, and task specific baselines on general PPI under our evaluation, and is the only method above chance on PPB-Affinity; on TCR-pMHC it leads at a fixed decision threshold but is outranked by a task specific sequence model. Because each protein is encoded independently in the first phase, its representation can be precomputed and reused across candidate pairs, which is convenient for large scale screening.