🤖 AI Summary
To address the challenge of precise literature retrieval for n-ary relation completion in biomedical knowledge bases, this paper proposes a knowledge-base-guided neural retrieval method. Specifically, it designs a hierarchical contrastive loss function to robustly model noisy and incomplete relational structures, targeting both missing argument completion and contextual evidence support; it further introduces a knowledge-base-informed balanced negative sampling strategy to enable efficient weakly supervised training. The method integrates neural retrieval, contrastive learning, and knowledge-base priors. On two biomedical retrieval benchmarks, it achieves state-of-the-art performance, improving NDCG@10 by 5.7 and 3.7 percentage points, respectively. Its core innovation lies in explicitly incorporating knowledge-base structural information into both the loss function design and negative sample construction—thereby significantly enhancing semantic matching capability for complex, multi-entity relations.
📝 Abstract
Curation of biomedical knowledge bases (KBs) relies on extracting accurate multi-entity relational facts from the literature - a process that remains largely manual and expert-driven. An essential step in this workflow is retrieving documents that can support or complete partially observed n-ary relations. We present a neural retrieval model designed to assist KB curation by identifying documents that help fill in missing relation arguments and provide relevant contextual evidence. To reduce dependence on scarce gold-standard training data, we exploit existing KB records to construct weakly supervised training sets. Our approach introduces two key technical contributions: (i) a layered contrastive loss that enables learning from noisy and incomplete relational structures, and (ii) a balanced sampling strategy that generates high-quality negatives from diverse KB records. On two biomedical retrieval benchmarks, our approach achieves state-of-the-art performance, outperforming strong baselines in NDCG@10 by 5.7 and 3.7 percentage points, respectively.