It was great to go on the DeepMind podcast! Check it out for takes on what's up with interpretability, how well any of it works, when we can/can't just read the chain of thought, how it helps with safety, and more
Google DeepMind (@GoogleDeepMind)
A model’s chain of thought acts like a scratch pad, offering a window into its reasoning. 📝
On the latest episode of our podcast, host @fryrsquared sits down with @NeelNanda5 to explore interpretability – the science of reverse engineering how neural networks learn and think.
Timecodes:
00:00 Introduction
02:41 Motivation for interpretability research
04:01 Mechanistic interpretability
08:14 Chain of thought monitoring
18:14 Interpretability techniques
35:00 Auditing models for safety
48:53 What comes next for interpretability
Video
— https://nitter.net/GoogleDeepMind/status/2075620281515929716#m