MEL: Legal Spanish Language Model

📅 2025-01-27
📈 Citations: 0
✹ Influential: 0
📄 PDF
đŸ€– AI Summary
To address the limited terminological understanding of multilingual pretrained models (e.g., XLM-RoBERTa) in the doubly sparse setting of low-resource languages and specialized domains—specifically Spanish legal texts—this work introduces the first deeply adapted language model for Spanish legal language. Built upon the XLM-RoBERTa-large architecture, our model is the first domain- and language-aligned pretrained model specifically trained on authoritative Spanish legal corpora (BOE and parliamentary texts). It integrates rigorous legal text cleaning, fine-grained segmentation, and a joint optimization strategy combining domain-specific self-supervised pretraining and task-oriented fine-tuning. Evaluated on legal named entity recognition and clause classification, our model achieves an average F1-score improvement of over 12% relative to the XLM-RoBERTa-large baseline, substantially alleviating the semantic modeling bottleneck at the intersection of cross-lingual transfer and domain specialization. The model and preprocessing pipeline are publicly released.

Technology Category

Natural Language Processing: Safety and RobustnessMachine Learning: Large Multimodal Models (LMMs)Application Domains: Humanities & Computational Social Science

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Large pretrained models with web dataSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Legal texts, characterized by complex and specialized terminology, present a significant challenge for Language Models. Adding an underrepresented language, such as Spanish, to the mix makes it even more challenging. While pre-trained models like XLM-RoBERTa have shown capabilities in handling multilingual corpora, their performance on domain specific documents remains underexplored. This paper presents the development and evaluation of MEL, a legal language model based on XLM-RoBERTa-large, fine-tuned on legal documents such as BOE (Bolet'in Oficial del Estado, the Spanish oficial report of laws) and congress texts. We detail the data collection, processing, training, and evaluation processes. Evaluation benchmarks show a significant improvement over baseline models in understanding the legal Spanish language. We also present case studies demonstrating the model's application to new legal texts, highlighting its potential to perform top results over different NLP tasks.
Problem

Research questions and friction points this paper is trying to address.

Pre-trained Models
Domain-specific Language
Spanish Legal Text
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spanish legal language model
XLM-RoBERTa-large pre-trained model
legal document comprehension
🔎 Similar Papers
No similar papers found.
D
David Betancur SĂĄnchez
Instituto de IngenierĂ­a del Conocimiento (IIC)
N
Nuria Aldama GarcĂ­a
Instituto de IngenierĂ­a del Conocimiento (IIC)
Álvaro Barbero Jiménez
Álvaro Barbero Jiménez
Universidad AutĂłnoma de Madrid, Instituto de IngenierĂ­a del Conocimiento
Machine learningdeep learningSupport Vector Machinesconvex optimizationnatural language processing
M
Marta Guerrero Nieto
Instituto de IngenierĂ­a del Conocimiento (IIC)
P
Patricia MarsĂ  Morales
Instituto de IngenierĂ­a del Conocimiento (IIC)
N
NicolĂĄs Serrano Salas
Instituto de IngenierĂ­a del Conocimiento (IIC)
C
Carlos GarcĂ­a HernĂĄn
Instituto de IngenierĂ­a del Conocimiento (IIC)
P
Pablo Haya Coll
Instituto de IngenierĂ­a del Conocimiento (IIC)
E
Elena Montiel Ponsoda
Ontology Engineering Group, Universidad Politécnica de Madrid
P
Pablo Calleja Ibåñez
Ontology Engineering Group, Universidad Politécnica de Madrid