The paper she's talking about: arxiv.org/abs/2209.10652
Link
Toy Models of Superposition
Neural networks often pack many unrelated concepts into a single neuron - a puzzling phenomenon known as 'polysemanticity' which makes interpretability much more challenging. This paper provides a...
arxiv.org