3CEL: A corpus of legal Spanish contract clauses

📅 2025-01-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Spanish legal contract NLP has long suffered from the dual scarcity of high-quality annotated corpora and effective collaboration mechanisms between legal domain experts and NLP researchers. To address this, we introduce 3CEL—the first fine-grained, human-annotated corpus specifically designed for Spanish public procurement documents—comprising 373 real-world contracts and 4,782 entity annotations. We systematically define and annotate 19 legally salient elements (e.g., performance deadlines, liability clauses, payment terms), thereby filling a critical gap in high-quality, task-oriented resources for Spanish legal NLP. Annotation follows a rigorous legal expert–linguist collaborative framework, incorporating multiple validation rounds and stringent quality control. The fully open-sourced 3CEL corpus empirically demonstrates substantial improvements in legal information extraction model performance and cross-document generalization. It establishes a foundational resource for intelligent, automated legal review in Spanish-speaking jurisdictions.

Technology Category

Natural Language Processing: Information ExtractionApplication Domains: Humanities & Computational Social ScienceHumans and AI: Crowd Sourcing and Human Computation

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web data
📝 Abstract
Legal corpora for Natural Language Processing (NLP) are valuable and scarce resources in languages like Spanish due to two main reasons: data accessibility and legal expert knowledge availability. INESData 2024 is a European Union funded project lead by the Universidad Polit'ecnica de Madrid (UPM) and developed by Instituto de Ingenier'ia del Conocimiento (IIC) to create a series of state-of-the-art NLP resources applied to the legal/administrative domain in Spanish. The goal of this paper is to present the Corpus of Legal Spanish Contract Clauses (3CEL), which is a contract information extraction corpus developed within the framework of INESData 2024. 3CEL contains 373 manually annotated tenders using 19 defined categories (4 782 total tags) that identify key information for contract understanding and reviewing.
Problem

Research questions and friction points this paper is trying to address.

Spanish Legal Language
Natural Language Processing (NLP)
Resource Limitations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spanish Legal NLP
Annotated Contracts Database
INESData 2024 Project
💼 Related Jobs
No related jobs found.
N
Nuria Aldama García
Instituto de Ingeniería del Conocimiento
P
Patricia Marsa Morales
Instituto de Ingeniería del Conocimiento
D
David Betancur Sánchez
Instituto de Ingeniería del Conocimiento
Álvaro Barbero Jiménez
Álvaro Barbero Jiménez
Universidad Autónoma de Madrid, Instituto de Ingeniería del Conocimiento
Machine learningdeep learningSupport Vector Machinesconvex optimizationnatural language processing
M
Marta Guerrero Nieto
Instituto de Ingeniería del Conocimiento
P
Pablo Haya Coll
Instituto de Ingeniería del Conocimiento
P
Patricia Martín Chozas
Ontology Engineering Group, Universidad Politécnica de Madrid
E
Elena Montiel Ponsoda
Ontology Engineering Group, Universidad Politécnica de Madrid