Twitter/X

VLA-JEPA is a VLA agent variant that uses V-JEPA2 conditioning during training to…

Brief

VLA-JEPA is a VLA agent variant that uses V-JEPA2 conditioning during training to add a JEPA world-model objective for action-relevant dynamics and human-video pretraining. At inference the world model is removed, leaving a VLA with a Qwen backbone and action head; a demo fine-tuned on 13 examples ran real time on NVIDIA Robotics DGX Spark (2026-06-06).

Source evidence

VLA-JEPA just dropped in LeRobot 🤖

What makes this model special is that it does not just learn what action to take from a given observation, it also leverages a JEPA world model to learn action-relevant dynamics.

During training, the VLA leverages V-JEPA2 by conditioning its predictor. This clever trick adds a world modeling objective to the training, which also allows pretraining on human videos.
At inference, the world model is dropped entirely, keeping only a standard VLA architecture: Qwen backbone and action head.

The demo here was only fine-tuned on 13 examples, showing great pretraining capability and running in real time on @NVIDIARobotics DGX Spark!

VLA-JEPA is the first world model to be ported to LeRobot, and I feel like it won't be the last 🚀

@Thom_Wolf @ClementDelangue

Video