Twitter/X

Author @maximelabonne reports DeepSeek used their open-perfectblend dataset to…

Brief

Author @maximelabonne announces that DeepSeek trained its new DSpark drafter using their open-perfectblend dataset (published 2026-06-27). The dataset is an open-source reproduction of "The Perfect Blend" paper and provides more than 1 million diverse prompts spanning math, chat, and code, which the author is promoting again.

Why it matters

Author @maximelabonne reports DeepSeek used their open-perfectblend dataset to train the new DSpark drafter (tweet published 2026-06-27).

Key details

  • open-perfectblend is an open-source reproduction of "The Perfect Blend" paper and contains over 1 million diverse prompts across math, chat, and code.
Source evidence

Fun surprise: DeepSeek used my open-perfectblend dataset to train their new DSpark drafter

Time to promote it again! It's an open-source reproduction of "The Perfect Blend" paper.

If you ever need >1M diverse prompts in math, chat, and code, it does the job.