ColdDDI: Evaluating Knowledge Utilization in Cold-Start Drug-Drug Interaction Prediction

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether cold-start drug-drug interaction (DDI) prediction models genuinely leverage pharmacological mechanism evidence. To this end, we construct a diagnostic benchmark based on DrugBank and employ masking interventions alongside channel sensitivity analysis to evaluate the actual reliance of models on molecular and knowledge graph (KG) evidence across zero-, one-, and two-unseen-drug scenarios. Our findings reveal that KG-enhanced baselines fail to effectively exploit shared mediating factors. In contrast, fine-tuned large language models recover 89–93% of interactions when such factors are present but achieve only 40–62% in their absence, empirically demonstrating a strong dependence of predictive performance on mechanistic evidence.
📝 Abstract
Cold-start drug-drug interaction (DDI) prediction tests whether models can identify clinically significant interactions for drugs without training-time interaction history. Existing benchmarks mostly report aggregate edge-prediction scores, leaving a key evaluation question unanswered: when models receive molecular, textual, or knowledge-graph (KG) evidence, do they actually use the evidence that pharmacologically supports the interaction? We introduce ColdDDI, a reconstructible diagnostic benchmark built from DrugBank 5.1.13, with 1,900 approved small-molecule drugs and 565,731 positive DDI pairs. ColdDDI evaluates pairs with zero, one, or two unseen drugs. It also annotates each interaction by whether it changes drug exposure or drug effect, and by whether the biomedical knowledge graph contains shared enzymes, transporters, or targets that can plausibly mediate the interaction. These annotations separate evidence availability from predictive dependence. We evaluate eight conventional DDI methods and 13 LLMs; for open-weight LLMs, we test five prompt patterns and use masking, drug replacement, and channel-sensitivity metrics to probe knowledge utilization. ColdDDI exposes that, in the hardest split where both drugs are unseen, the main performance divide is mediator availability. A fine-tuned 1B LLM recovers 89-93% of interactions with a shared enzyme, transporter, or target, but only 40-62% without such a mediator. More importantly, KG-provided evidence is not always used; several KG-augmented baselines change little when the shared mediator is masked or disrupted, whereas fine-tuned LLMs respond strongly to this intervention. Thus, ColdDDI evaluates knowledge utilization rather than knowledge access alone, showing where cold-start DDI models rely on mechanistic evidence and where they fail despite receiving it. Code is available at https://github.com/0217ljh/ColdDDI-NeurIPS2026.
Problem

Research questions and friction points this paper is trying to address.

cold-start drug-drug interaction prediction
knowledge utilization
model evaluation
biomedical knowledge graph
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cold-start DDI prediction
Diagnostic benchmark
Knowledge utilization
Large language models
Biomedical knowledge graph
🔎 Similar Papers
No similar papers found.