Guidelines for the Creation of an Annotated Corpus

📅 2026-01-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic approaches for constructing, storing, and sharing high-quality annotated corpora. It proposes a generalizable and reusable end-to-end methodology encompassing annotation guideline development, corpus annotation, data storage, sharing mechanisms, and value realization, with an emphasis on full lifecycle management and cross-domain applicability. Integrating linguistic annotation theory, data management standards, and collaborative research practices, the approach is articulated through a structured framework and illustrative examples to yield a clear and actionable guide. The resulting methodology provides standardized support for diverse research domains, significantly enhancing the efficiency and quality with which researchers can build and utilize annotated textual data.

Technology Category

Natural Language Processing: Lexical Semantics and MorphologyApplication Domains: Humanities & Computational Social ScienceData Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Semantics and Knowledge: Representation, semantic annotation, enhancement, enrichments, access and/or integration of a variety of data on the WebEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web data
📝 Abstract
This document, based on feedback from UMR TETIS members and the scientific literature, provides a generic methodology for creating annotation guidelines and annotated textual datasets (corpora). It covers methodological aspects, as well as storage, sharing, and valorization of the data. It includes definitions and examples to clearly illustrate each step of the process, thus providing a comprehensive framework to support the creation and use of corpora in various research contexts.
Problem

Research questions and friction points this paper is trying to address.

annotated corpus
annotation guidelines
textual datasets
data sharing
corpus creation
Innovation

Methods, ideas, or system contributions that make the work stand out.

annotation guidelines
annotated corpus
textual dataset
corpus methodology
data valorization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Bahdja Boudoua
TETIS, Univ. Montpellier, AgroParisTech, CIRAD, CNRS, INRAE, Montpellier, France; INRAE, UMR TETIS, Montpellier, France
N
Nadia Guiffant
TETIS, Univ. Montpellier, AgroParisTech, CIRAD, CNRS, INRAE, Montpellier, France; INRAE, UMR TETIS, Montpellier, France
M
M. Roche
TETIS, Univ. Montpellier, AgroParisTech, CIRAD, CNRS, INRAE, Montpellier, France; CIRAD, UMR TETIS, F-34398 Montpellier, France
M
M. Teisseire
TETIS, Univ. Montpellier, AgroParisTech, CIRAD, CNRS, INRAE, Montpellier, France; INRAE, UMR TETIS, Montpellier, France
A
Annelise Tran
TETIS, Univ. Montpellier, AgroParisTech, CIRAD, CNRS, INRAE, Montpellier, France; INRAE, UMR TETIS, Montpellier, France; UMR ASTRE, Univ. Montpellier, CIRAD, INRAE, Montpellier, France