Twitter/X

Rohit Girdhar announced on 2026-06-16 that he is leaving Meta after almost 7…

Brief

Rohit Girdhar announced on 2026-06-16 that, after almost 7 years at Meta, he’s joining AMI Labs in New York to build real-world models. He emphasizes video-temporal research—predicting future visual states—and cites contributions to Ego4D, Omnivore, ImageBind, Emu Video, Llama-3 and MovieGen as groundwork for extending AI beyond language-model capabilities into physical, multimodal tasks.

Why it matters

Rohit Girdhar announced on 2026-06-16 that he is leaving Meta after almost 7 years to join AMI Labs in New York to “help build real-world models.”

Key details

  • His research focuses on leveraging the temporal structure of video to predict future visual states (analogous to next-token language models), advancing physical reasoning, action anticipation, image animation, and video generation; he worked on Ego4D, Omnivore, ImageBind, Emu Video, Llama-3, and MovieGen.
  • He frames AMI’s mission as the next frontier after language-model successes in coding, math, and reasoning, and thanks collaborators @sainingxie, @michaelrabbat, Min Lin, @jingli9111, @ylecun, @lxbrun and the AMI team.
Source evidence

After almost 7 incredible years at Meta, I’m excited to share that I’ll be joining the team at AMI Labs in New York to help build real-world models!

AMI’s mission is something that has been close to my heart since my early days at FAIR. Much of my work over the years has focused on building perception and generation models that leverage the temporal structure of videos to predict future visual states, much like how language models leverage the structure of language to predict the next token. Along the way, this work helped advance a range of visual tasks, including physical reasoning, action anticipation, image animation, and video generation, while also contributing to some of the largest video datasets and benchmarks in the world.

Now that language models have demonstrated remarkable capabilities in coding, math, and reasoning, I believe it’s the right time to push toward the next frontier: models that can achieve similar impact across a much broader range of real-world tasks.

Leaving Meta has been one of the hardest decisions of my life. It was my first job after grad school, and in many ways, I grew up there. Many of the people I met there became close friends, mentors, and collaborators. When I joined, billion-parameter models were just beginning to emerge, vision systems were largely built separately for each task, and video generation models were hard to control and could produce little more than low-resolution clips. Needless to say, a lot changed over the years, and I was incredibly fortunate to have a front-row seat to that progress and to contribute along the way.

I had the privilege of working on projects spanning multimodal modeling, representation learning, and video generation, including Ego4D, Omnivore, ImageBind, Emu Video, Llama-3, and MovieGen. None of this would have been possible without the exceptionally talented and kind collaborators I had the opportunity to work alongside. I’ll always be thankful to Meta for the opportunities, growth, learning, and most importantly, the friendships that will stay with me for a lifetime.

Looking ahead, I’m deeply grateful to @sainingxie @michaelrabbat Min Lin @jingli9111 @ylecun @lxbrun and the entire AMI team for the opportunity to join this ambitious mission. I’m excited to help build the next generation of real-world models and push toward AI systems that better understand and interact with the physical world.

Onwards! 🚀