Twitter/X

VERA is a 14-billion-parameter video-to-action system from MIT (announced…

Brief

VERA, announced by MIT on 2026-06-23, is a 14B video-to-action model that converts video world models into embodiment-agnostic robot policies. The team demonstrates one set of weights controlling diverse hardware and skills — including zero-shot pick-and-place on a Franka Panda arm and contact-rich manipulation with a 16-DoF hand — and is releasing the code and models at vera.csail.mit.edu for fine-tuning.

Why it matters

VERA is a 14-billion-parameter video-to-action system from MIT (announced 2026-06-23) designed to produce embodiment-agnostic robot policies.

Key details

  • VERA runs the same video planner with the same weights across different robots and tasks — demonstrated zero-shot pick-and-place on a real Franka Panda arm and contact-rich cube reorientation with a 16-DoF robotic hand.
  • The project is being open-sourced (vera.csail.mit.edu) so users can fine-tune VERA for their own robot setups and environments.
Source evidence

Robot learning is moving beyond policies built for one robot, one scene, one task.

At MIT, we’re exploring a different path: turning video world models into embodiment-agnostic robot policies.

Introducing VERA: a 14B video-to-action system that controls robots across embodiments, skills, and environments.

From zero-shot pick-and-place on a real Panda arm to contact-rich cube reorientation with a 16-DoF robotic hand.

Different robots. Different environments. Different tasks.
Same video planner. Same weights.

We’re open-sourcing everything so you can fine-tune VERA for your own robot setup too. Deep dive in the thread:

🔗 vera.csail.mit.edu
🧵 (1/7)

Video