Twitter/X

Model 'pi-0.5' was trained on roughly 40,000 hours of data according to the post.

Brief

The post compares dataset sizes for two models—pi-0.5 (~40,000 training hours) versus dreamzero (~500 hours)—to argue that even a few hundred to a few tens of thousands of hours can be sufficient to start training. It points readers to datasets.bot, which the author and a commenter say aggregates open-source datasets useful for model development.

Why it matters

Model 'pi-0.5' was trained on roughly 40,000 hours of data according to the post.

Key details

  • 'dreamzero' was trained on about 500 hours, implying a much smaller dataset.
  • The author recommends datasets.bot as an aggregator of open-source datasets for getting started.
Source evidence

Aggregator of open-source datasets ->

pi-0.5 was ~40k hours
dreamzero only 500

you should have enough to get started with

Vai Viswanathan (@vai_viswanathan)

check out datasets.bot! Been aggregating these!

— https://nitter.net/vai_viswanathan/status/2070011384289460335#m