ArXiv

FRENCH-YMCA: A FRENCH Corpus meeting the language needs of Youth, froM Children to Adolescents

Authors
Cherifa Ben Khelil, Jean-Yves Antoine, Anaïs Halftermeyer...
Categories
cs.CL
arXiv
https://arxiv.org/abs/2604.05899v1
PDF
https://arxiv.org/pdf/2604.05899v1

Brief

The French-YMCA corpus is a new linguistic resource targeting children and adolescents, comprising 39,200 texts and 22,471,898 words drawn from diverse sources with consistent grammar and spelling. Designed for open online access, it aims to support development of language models that understand youth-specific language and produce age-appropriate interactions; full details are on the 2026-04-07 arXiv submission.

Source evidence

title: FRENCH-YMCA: A FRENCH Corpus meeting the language needs of Youth, froM Children to Adolescents
author: Cherifa Ben Khelil, Jean-Yves Antoine, Anaïs Halftermeyer...
contenttype: arxivpaper
publication: ArXiv
published: 2026-04-07T14:04:30+00:00
source_url: https://arxiv.org/abs/2604.05899v1

word_count: 151

Abstract

In this paper, we introduce the French-YMCA corpus, a new linguistic resource specifically tailored for children and adolescents. The motivation for building this corpus is clear: children have unique language requirements, as their language skills are in constant evolution and differ from those of adults. With an extensive collection of 39,200 text files, the French-YMCA corpus encompasses a total of 22,471,898 words. It distinguishes itself through its diverse sources, consistent grammar and spelling, and the commitment to providing open online accessibility for all. Such corpus can serve as the foundation for training language models that understand and anticipate youth's language, thereby enhancing the quality of digital interactions and ensuring that responses and suggestions are age-appropriate and adapted to the comprehension level of users of this age.

Comment: 5 pages, 1 figure