YouTube

They Looked Inside Claude’s AI's Mind. It Got Weird

Brief

Two Minute Papers (tutorial-style video, 2026-06-16) summarizes recent work on Natural Language Autoencoders found inside Claude, describing how researchers probed internal activations and discovered compressed, interpretable representations and surprising internal behaviors. The segment links to Anthropic’s paper and a Transformer Circuits walkthrough for technical details and visualizations.

Why it matters

Two Minute Papers (video published 2026-06-16) reviews research revealing natural-language autoencoder (NLA)–like structures inside Anthropic’s Claude, showing compressed, interpretable internal representations in model layers.

Key details

  • The video cites two primary sources: Anthropic’s 'Natural Language Autoencoders' paper and a Transformer Circuits explainer (2026 NLA), linked at https://www.anthropic.com/research/natural-language-autoencoders and https://transformer-circuits.pub/2026/nla/index.html.
Source evidence

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers

📝 The paper is available here:
https://www.anthropic.com/research/natural-language-autoencoders
https://transformer-circuits.pub/2026/nla/index.html

🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi

My research: https://cg.tuwien.ac.at/~zsolnai/
Thumbnail design: https://felicia.hu

Channel: Two Minute Papers
Published: 2026-06-16
Video URL: https://www.youtube.com/watch?v=l72ufA-4SzE