Twitter/X

Anthropic's 2022 paper 'Toy Models of Superposition' shows models encode concepts…

Brief

How an AI model holds the full weight of a word like 'grief' is explained using Anthropic's 2022 paper Toy Models of Superposition: concepts are stored as directions in activation space rather than single neurons, producing neural 'superposition' where individual neurons mix fragments of many concepts, so no exact locus for 'grief'—see Drishti Grover's plain-language video.

Why it matters

Anthropic's 2022 paper 'Toy Models of Superposition' shows models encode concepts as directions in activation space rather than dedicating one neuron per concept, allowing far more concepts than neurons.

Key details

  • Individual neurons are in 'superposition' and carry fragments of many unrelated ideas, so there is no exact spot where an AI 'understands' grief—the meaning is distributed and entangled across activations.
  • @alex_prompter (2026-08-06) recommends Drishti Grover's plain-language video as the clearest walkthrough of this idea.
Source evidence

How does an AI model hold the full weight of a word like "grief"?

Anthropic dug into this back in 2022 with a paper called Toy Models of Superposition. The short version is that models don't keep one neuron per concept. They pack far more concepts than they have neurons by storing each one as a direction in space, so a single neuron ends up carrying pieces of many unrelated ideas at once.

That's why nobody can point to the exact spot where an AI understands grief. The meaning is spread out and tangled with the rest of what the model knows.

Drishti Grover explains the whole idea in plain language in this video. It's the clearest walkthrough of the concept I've come across, and worth your time.

Video