ClinText-SP and RigoBERTa Clinical: a new set of open resources for Spanish Clinical NLP

📅 2025-03-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Spanish clinical NLP faces dual bottlenecks: scarcity of publicly available annotated corpora and absence of dedicated language models. To address this, we introduce ClinText-SP—the first large-scale, high-quality, open-source Spanish clinical corpus, comprising de-identified multi-source clinical texts. Building upon it, we propose RigoBERTa Clinical, a domain-adapted language model trained via clinical-domain self-supervised pretraining. Our methodology integrates cross-institutional clinical text cleaning, multi-task annotation alignment, and injection of clinical domain knowledge. Experiments demonstrate that RigoBERTa Clinical achieves state-of-the-art performance on Spanish clinical NER, question answering, and text classification benchmarks, significantly outperforming existing general-purpose and multilingual models. ClinText-SP has already enabled multiple downstream studies and gained broad adoption within the international research community. This work fills a critical gap in foundational resources for Spanish clinical NLP and establishes a reusable paradigm for developing medical AI in low-resource languages.

Technology Category

Natural Language Processing: Language Grounding & Multi-modal NLPMachine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Large pretrained models with web dataEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasets
📝 Abstract
We present a novel contribution to Spanish clinical natural language processing by introducing the largest publicly available clinical corpus, ClinText-SP, along with a state-of-the-art clinical encoder language model, RigoBERTa Clinical. Our corpus was meticulously curated from diverse open sources, including clinical cases from medical journals and annotated corpora from shared tasks, providing a rich and diverse dataset that was previously difficult to access. RigoBERTa Clinical, developed through domain-adaptive pretraining on this comprehensive dataset, significantly outperforms existing models on multiple clinical NLP benchmarks. By publicly releasing both the dataset and the model, we aim to empower the research community with robust resources that can drive further advancements in clinical NLP and ultimately contribute to improved healthcare applications.
Problem

Research questions and friction points this paper is trying to address.

Lack of large publicly available Spanish clinical NLP corpus
Need for state-of-the-art clinical encoder language model in Spanish
Limited resources for advancing clinical NLP in healthcare applications
Innovation

Methods, ideas, or system contributions that make the work stand out.

Largest Spanish clinical corpus ClinText-SP
State-of-the-art model RigoBERTa Clinical
Domain-adaptive pretraining on diverse data
🔎 Similar Papers
No similar papers found.