๐ค AI Summary
Existing dense retrievers exhibit poor out-of-distribution generalization in clinical settings and lack multilingual evaluation benchmarks tailored to medical decision-making. To address this, we introduce CUREโthe first point-of-careโoriented, multilingual clinical retrieval dataset. CURE spans 10 medical domains and comprises 2,000 real-world clinical queries, supporting both monolingual (English) and cross-lingual (French/Spanish โ English) evaluation. It was co-constructed by clinical experts using an ad-hoc paradigm combining manual annotation with expert verification, and is compatible with both dense and sparse retrieval models under standard IR metrics (MRR, NDCG@10). The dataset and its open-source implementation are integrated into the Hugging Face ecosystem. Empirical evaluation reveals substantial performance degradation of state-of-the-art dense retrievers on CURE, exposing critical limitations in clinical domain generalization. CURE has emerged as a new de facto standard for evaluating medical AI retrieval systems.
๐ Abstract
Given the dominance of dense retrievers that do not generalize well beyond their training dataset distributions, domain-specific test sets are essential in evaluating retrieval. There are few test datasets for retrieval systems intended for use by healthcare providers in a point-of-care setting. To fill this gap we have collaborated with medical professionals to create CURE, an ad-hoc retrieval test dataset for passage ranking with 2000 queries spanning 10 medical domains with a monolingual (English) and two cross-lingual (French/Spanish ->English) conditions. In this paper, we describe how CURE was constructed and provide baseline results to showcase its effectiveness as an evaluation tool. CURE is published with a Creative Commons Attribution Non Commercial 4.0 license and can be accessed on Hugging Face.