FRENCH-YMCA: A FRENCH Corpus meeting the language needs of Youth, froM Children to Adolescents

📅 2026-04-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Children and adolescents are in a critical period of language development, exhibiting linguistic characteristics that markedly differ from those of adults; however, age-appropriate, large-scale, and high-quality French textual resources remain scarce. To address this gap, this study presents French-YMCA, the first systematically constructed open-access corpus spanning the full developmental range from childhood through adolescence. Comprising 39,200 normative texts totaling over 22.47 million words, the corpus was assembled through multi-source collection, rigorous cleaning, and standardized processing. An accompanying open-access platform ensures broad usability. French-YMCA fills a critical void in French youth language resources and provides essential support for training age-adapted language models and enhancing the comprehensibility and developmental appropriateness of digital interactions.

Technology Category

Natural Language Processing: Lexical Semantics and MorphologyApplication Domains: Humanities & Computational Social ScienceMachine Learning: Large Multimodal Models (LMMs)

Application Category

Web Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSemantics and Knowledge: Scalable techniques for the creation, curation, publication, maintenance, and consumption of large, Web-based, structured, reusable, knowledge graphs and ontologies
📝 Abstract
In this paper, we introduce the French-YMCA corpus, a new linguistic resource specifically tailored for children and adolescents. The motivation for building this corpus is clear: children have unique language requirements, as their language skills are in constant evolution and differ from those of adults. With an extensive collection of 39,200 text files, the French-YMCA corpus encompasses a total of 22,471,898 words. It distinguishes itself through its diverse sources, consistent grammar and spelling, and the commitment to providing open online accessibility for all. Such corpus can serve as the foundation for training language models that understand and anticipate youth's language, thereby enhancing the quality of digital interactions and ensuring that responses and suggestions are age-appropriate and adapted to the comprehension level of users of this age.
Problem

Research questions and friction points this paper is trying to address.

child language
adolescent language
language resources
age-appropriate interaction
youth linguistic needs
Innovation

Methods, ideas, or system contributions that make the work stand out.

French-YMCA corpus
child language modeling
age-appropriate NLP
open linguistic resource
youth language development
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Cherifa Ben Khelil
EFREI Research Lab - University of Paris Panthéon Assas, France
J
Jean-Yves Antoine
LIFAT - University of Tours, France
A
Anaïs Halftermeyer
LIFO - University of Orleans, France
F
Frédéric Rayar
LIFAT - University of Tours, France
M
Mathieu Thebaud
CMRRF, Mutuality of Morbihan, France