Twitter/X

Nemotron-Personas-Belgium released on 2026-06-17 in collaboration with NVIDIA AI…

Brief

Nemotron-Personas-Belgium is a synthetic dataset published 2026-06-17 (co-developed with NVIDIA AI) that comprises 1.2 million Belgian personas across four language variants (Dutch, French, German, English). The dataset mirrors the Statbel census distributions and includes demographic and socioeconomic attributes such as age, occupation, household structure, income, and names.

Why it matters

Nemotron-Personas-Belgium released on 2026-06-17 in collaboration with NVIDIA AI, providing synthetic persona data for Belgium.

Key details

  • Dataset contains 4 × 300,000 (1.2 million) synthetic Belgian personas in Dutch, French, German, and English, modeled on the Statbel census with attributes including age, occupations, household composition, income, and names.
Source evidence

And more details from @pieterdelobelle

Pieter Delobelle (@pieterdelobelle)

Today we release Nemotron-Personas-Belgium in collaboration with @NVIDIAAI , a synthetic dataset of 4x300k Belgian personas in Dutch, French, German and English.

The dataset is modeled on the @Statbel_en census, modeling age, occupations, household, income, names etc.

— https://nitter.net/pieterdelobelle/status/2067275709978915059#m