And new data release: in partnership with @GSMA we publish Telco-Common-Corpus, the largest fully open corpus for the telecom sector, 10 billion tokens all under free license/public domain.
@Dorialexander announced release of Telco-Common-Corpus in partnership with GSMA
Brief
Telco-Common-Corpus, announced by @Dorialexander on 2026-06-25 in partnership with GSMA, is a 10-billion-token dataset positioned as the largest fully open corpus for telecommunications. The release makes the data available under a free/public-domain license, enabling researchers and developers to use it without proprietary restrictions for telecom-focused NLP and LLM training.
Why it matters
@Dorialexander announced release of Telco-Common-Corpus in partnership with GSMA
Key details
- Corpus size: 10 billion tokens, described as the largest fully open corpus for the telecom sector
- License: all data released under a free license / public domain