🤖 AI Summary
Spanish legal contract NLP has long suffered from the dual scarcity of high-quality annotated corpora and effective collaboration mechanisms between legal domain experts and NLP researchers. To address this, we introduce 3CEL—the first fine-grained, human-annotated corpus specifically designed for Spanish public procurement documents—comprising 373 real-world contracts and 4,782 entity annotations. We systematically define and annotate 19 legally salient elements (e.g., performance deadlines, liability clauses, payment terms), thereby filling a critical gap in high-quality, task-oriented resources for Spanish legal NLP. Annotation follows a rigorous legal expert–linguist collaborative framework, incorporating multiple validation rounds and stringent quality control. The fully open-sourced 3CEL corpus empirically demonstrates substantial improvements in legal information extraction model performance and cross-document generalization. It establishes a foundational resource for intelligent, automated legal review in Spanish-speaking jurisdictions.
📝 Abstract
Legal corpora for Natural Language Processing (NLP) are valuable and scarce resources in languages like Spanish due to two main reasons: data accessibility and legal expert knowledge availability. INESData 2024 is a European Union funded project lead by the Universidad Polit'ecnica de Madrid (UPM) and developed by Instituto de Ingenier'ia del Conocimiento (IIC) to create a series of state-of-the-art NLP resources applied to the legal/administrative domain in Spanish. The goal of this paper is to present the Corpus of Legal Spanish Contract Clauses (3CEL), which is a contract information extraction corpus developed within the framework of INESData 2024. 3CEL contains 373 manually annotated tenders using 19 defined categories (4 782 total tags) that identify key information for contract understanding and reviewing.