TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance overestimation in existing ligand virtual screening benchmarks caused by random negative sampling by constructing a multi-target evaluation benchmark based on hard negatives. The proposed benchmark encompasses 93 protein targets and employs structurally similar decoy libraries with a fixed 1:40 ratio to eliminate physicochemical property shortcuts, alongside three standardized evaluation protocols. Through systematic comparisons integrating ChEMBL data with molecular fingerprints, graph neural networks, and state-of-the-art molecular representation models, this work reveals that the performance of existing methods degrades substantially when evaluated against structurally similar decoys. All code and datasets have been open-sourced, providing a reproducible and equitable comparison platform for ligand-based virtual screening (LBVS) methodologies.
📝 Abstract
Ligand-based virtual screening (LBVS) is a practical first-pass tool in early-stage drug discovery, but existing benchmarks can overestimate performance through random negatives, easy decoys, limited target coverage, and non-standardized evaluation protocols. We introduce TopU-LBVS, a multi-target benchmark for LBVS under hard-negative screening conditions. Starting from curated ChEMBL~35 bioactivity data, TopU-LBVS covers 93 protein targets across 7 protein classes and constructs target-specific screening libraries with property-matched, structurally similar decoys at a fixed 1:40 active-to-decoy ratio. Libraries contain roughly 400 to 10,000 compounds and are designed to reduce simple physicochemical and nearest-neighbor fingerprint shortcuts. TopU-LBVS provides three fixed protocols. TopU-LBVS-full evaluates ChEMBL$^\ast \rightarrow$ TopU generalization across all 93 targets. TopU-LBVS-low evaluates low-data TopU $\rightarrow$ TopU learning within the hard-negative distribution. TopU-LBVS-mini provides a compact seven-target protocol with a paired random-decoy control that changes only the test decoys, enabling low-cost development and direct measurement of the gap between random ChEMBL$^\ast$ and TopU decoys. Across ten reference baselines spanning fingerprint methods, molecular GNNs, fingerprint hybrids, and modern molecular models, performance under random-decoy evaluation degrades sharply under hard-negative screening. We release data, fixed splits, evaluation code, and baseline implementations for reproducible comparison of future LBVS and molecular representation learning methods. Code and data are available at https://github.com/topu-benchmark/topu-lbvs and https://huggingface.co/datasets/topu-benchmark/topu-lbvs.
Problem

Research questions and friction points this paper is trying to address.

Ligand-based virtual screening
benchmark
hard negatives
decoys
drug discovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ligand-based virtual screening
Hard-negative benchmark
Multi-target evaluation
Molecular representation learning
Decoy design
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Surbhi Kumar
UT Dallas, Math. Sci.
Y
Yuhe Zhou
National Inst. Bio. Sci.
V
Varun Shiralkar
UT Dallas, Comp. Sci.
N
Niu Huang
National Institute of Biological Sciences
Baris Coskunuzer
Baris Coskunuzer
University of Texas at Dallas
Geometric TopologyTopological Data AnalysisMachine Learning