BERT4DTI : BERT-based Model for Predicting Drug-Protein Interactions

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of scarce labeled data, high fine-tuning costs, and the lack of interactive dependencies in independent sequence encoding for drug-target prediction by proposing an efficient predictive framework. The method employs ChemBERTa and ProtBERT as encoders and introduces a bidirectional cross-attention mechanism to capture specific dependencies between drug-protein pairs, followed by convolutional layers and a multilayer perceptron for classification. Additionally, a truncated layer strategy is adopted to substantially reduce the number of trainable parameters. Experimental results demonstrate that the proposed model achieves state-of-the-art ROC-AUC and sensitivity across multiple benchmarks while reducing the total parameter count to 125 million, thereby realizing an excellent balance between predictive performance and computational efficiency.
📝 Abstract
Understanding how drugs interact with protein targets is fundamental to drug discovery, drug repurposing and the early identification of promising therapeutic candidates before costly experimental testing. Sequence-based DTI models face three practical limitations: labelled interactions are scarce and unevenly distributed, large pretrained chemical and protein encoders are expensive to fine-tune end-to-end, and independently encoded sequences do not capture pair-specific dependencies. We present BERT4DTI, which encodes SMILES strings with ChemBERTa and amino-acid sequences with ProtBERT, applies bidirectional mutual attention between token-level representations, and classifies the resulting interaction features using convolutional layers and a multilayer perceptron. To reduce trainable size, ProtBERT is truncated to 18 retained layers and only the last two layers of each encoder are fine-tuned. On BIOSNAP, DAVIS and BindingDB, BERT4DTI is competitive, achieving the best ROC-AUC and PR-AUC on BIOSNAP and the highest sensitivity on all three benchmarks. An ablation on DAVIS shows that mutual attention improves PR-AUC and specificity. With 125M trainable parameters compared with 353M for full BERT fine-tuning, BERT4DTI provides a favourable performance-parameter trade-off for sequence-based DTI screening, while leaving runtime profiling, calibration and leakage-audited validation for future work.
Problem

Research questions and friction points this paper is trying to address.

Drug-Target Interaction
Sequence-based prediction
Data scarcity
Fine-tuning cost
Pair-specific dependencies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Drug-Protein Interaction
Bidirectional Mutual Attention
Parameter-Efficient Fine-Tuning
ChemBERTa
ProtBERT
💼 Related Jobs
No related jobs found.
T
Thanina Hamitouch
Ecole Nationale Supérieure d’Informatique (ESI), Algiers, Algeria
K
Khadidja Henni
Institut d’Intelligence Artificielle Appliquée, TELUQ University, Montreal, Canada
A
Abdelkrim Arie
Ecole Nationale Supérieure d’Informatique (ESI), Algiers, Algeria
A
Amina Selma Haichour
Ecole Nationale Supérieure d’Informatique (ESI), Algiers, Algeria
Neila Mezghani
Neila Mezghani
Institut d’Intelligence Artificielle Appliquée, TELUQ University, Montreal, Canada
L
Lina Abou-Abbas
Department of Electrical and Computer Engineering, Lebanese American University, Byblos, Lebanon