Addressing Model Overcomplexity in Drug-Drug Interaction Prediction With Molecular Fingerprints

📅 2025-03-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address high model complexity, poor generalizability, and excessive computational cost in drug–drug interaction (DDI) prediction, this work proposes a lightweight neural network–based minimalist modeling paradigm. Methodologically, we systematically evaluate three molecular representations—Morgan fingerprints, GCN-based graph embeddings, and MoLFormer molecular transformer embeddings—and construct compact fully connected models under leak-proof data splitting. We further integrate Grad-CAM–style gradient analysis for interpretability validation. Key contributions include: (1) demonstrating that simple molecular representations achieve state-of-the-art performance (AUC > 0.92 on DrugBank/FDA datasets) even under stringent generalization settings; (2) the first identification—via interpretable analysis—of clinically relevant pharmacophores, including CYP inhibition and P-glycoprotein substrate motifs; and (3) challenging the “complexity bias” in DDI modeling by advocating a new paradigm prioritizing data quality and incremental model development over architectural sophistication.

Technology Category

Knowledge Representation and Reasoning: Diagnosis and Abductive ReasoningMachine Learning: Deep Neural Architectures and Foundation ModelsData Mining & Knowledge Management: Other Foundations of Data Mining & Knowledge Management

Application Category

Graph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsUser Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Accurately predicting drug-drug interactions (DDIs) is crucial for pharmaceutical research and clinical safety. Recent deep learning models often suffer from high computational costs and limited generalization across datasets. In this study, we investigate a simpler yet effective approach using molecular representations such as Morgan fingerprints (MFPS), graph-based embeddings from graph convolutional networks (GCNs), and transformer-derived embeddings from MoLFormer integrated into a straightforward neural network. We benchmark our implementation on DrugBank DDI splits and a drug-drug affinity (DDA) dataset from the Food and Drug Administration. MFPS along with MoLFormer and GCN representations achieve competitive performance across tasks, even in the more challenging leak-proof split, highlighting the sufficiency of simple molecular representations. Moreover, we are able to identify key molecular motifs and structural patterns relevant to drug interactions via gradient-based analyses using the representations under study. Despite these results, dataset limitations such as insufficient chemical diversity, limited dataset size, and inconsistent labeling impact robust evaluation and challenge the need for more complex approaches. Our work provides a meaningful baseline and emphasizes the need for better dataset curation and progressive complexity scaling.
Problem

Research questions and friction points this paper is trying to address.

Reducing computational costs in drug interaction prediction
Improving generalization across diverse DDI datasets
Identifying key molecular motifs for interactions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses Morgan fingerprints for molecular representation
Integrates GCN and MoLFormer embeddings
Employs gradient-based motif analysis
💼 Related Jobs
No related jobs found.
M
Manel Gil-Sorribes
Nostrum Biodiscovery, Barcelona, 08029, Spain
A
Alexis Molina
Nostrum Biodiscovery, Barcelona, 08029, Spain