This is really cool work building off our “Interpreting Physics in Video World Models” paper. People may not realize this yet, but unsupervised latent space simulators are a big deal.
Jay Hack (@mathemagic1an)
Do video models actually learn physics? Or are they just stochastic pixel parrots?
If you train a transformer on raw pixels, you find a causal world model - and you can even "play" it like a video game
This is Anthropic's J-lens applied to video models
Code + game below 👇
— https://nitter.net/mathemagic1an/status/2082482976957558836#m