How does an AI model hold the full weight of a word like "grief"?
Anthropic dug into this back in 2022 with a paper called Toy Models of Superposition. The short version is that models don't keep one neuron per concept. They pack far more concepts than they have neurons by storing each one as a direction in space, so a single neuron ends up carrying pieces of many unrelated ideas at once.
That's why nobody can point to the exact spot where an AI understands grief. The meaning is spread out and tangled with the rest of what the model knows.
Drishti Grover explains the whole idea in plain language in this video. It's the clearest walkthrough of the concept I've come across, and worth your time.
Video