Twitter/X

@cryptopunk7213 (published 2026-08-06) claims Google "basically owns the richest…

Brief

@cryptopunk7213 (published 2026-08-06) argues Google "basically owns the richest corpus" of training data—"every site you visit, every click, every thing you search"—and reports that Meta is allegedly building its own web index/search engine (via a @levelsio DM) to capture intent across WhatsApp, Instagram and Facebook; combined with Meta's compute, this could make Meta competitive with Google.

Why it matters

@cryptopunk7213 (published 2026-08-06) claims Google "basically owns the richest corpus" for training frontier models because "every site you visit, every click you make, every thing you search" is documented and used to predict user behavior.

Key details

  • Meta is ALLEGEDLY building its own web index/search engine (reported via a DM to @levelsio) to capture user intent from WhatsApp, Instagram and Facebook so AI web searches don't route to Google and feed Google's training.
  • The author asserts that pairing Meta's claimed intent data capture with Meta's compute resources could make Meta "pretty competitive" with Google in AI.
Source evidence

this is nuts. google basically owns the richest corpus of data that can be used to train frontier models:

  • every site you visit, every click you make, every thing you search is all documented, filed and used to create products that predict your next step

  • ai models are the ultimate form factor for this and meta knows it.

  • so metas allegedly creating their own damn search engine to capture everyone’s intent and actions.

they own user intent data on social media across whatsapp, instagram and facebook so this is their hail mary to capture google’s share

pair this with their compute and you can see how meta could become pretty competitive

@levelsio (@levelsio)

Meta staff DM'd me secretly

Posted with permission

Meta is ALLEGEDLY building their own Google search engine, so that if their AI does a web search it doesn't end up at Google, as Google could then use it for THEIR training, so they want their own web index that they will then use as their own Meta search engine for their AI

Interesting 🤔

— https://nitter.net/levelsio/status/2085467097405247945#m